python-resilience
Python resilience patterns including automatic retries, exponential backoff, timeouts, and fault-tolerant decorators. Use when adding retry logic, implementing timeouts, building fault-tolerant services, or handling transient failures.
By wshobson · 10,197 installs
npx skills add wshobson/agents --skill python-resilience
Source repository · Upstream listing
Python Resilience Patterns
Build fault tolerant Python applications that gracefully handle transient failures, network issues, and service outages. Resilience patterns keep systems running when dependencies are unreliable.
When to Use This Skill
Adding retry logic to external service calls
Implementing timeouts for network operations
Building fault tolerant microservices
Handling rate limiting and backpressure
Creating infrastructure decorators
Designing circuit breakers
Core Concepts
1. Transient vs Permanent Failures
Retry transient errors (network timeouts, temporary service issues). Don't retry permanent errors (invalid credentials, bad requests).
2. Exponential Backoff
Increase wait time between retries to avoid overwhelming recovering services.
3. Jitter
Add randomness to backoff to prevent thundering herd when many clients retry simultaneously.
4. Bounded Retries
Cap both attempt count and total duration to prevent infinite retry loops.
Quick Start
Fundamental Patterns
Pattern 1: Basic Retry with Tenacity
Use the tenacity library for production grade retry logic. For simpler cases, consider built in retry functionality or a lightweight custom implementation.
Pattern 2: Retry Only Appropriate Errors
Whitelist specific transient exceptions. Never retry:
ValueError , TypeError These are bugs, not transient issues
AuthenticationError Invalid credentials won't become valid
HTTP 4xx errors (except 429) Client errors are permanent
Pattern 3: HTTP Status Code Retries
Retry specific HTTP status codes that indicate transient issues.
Pattern 4: Combined Exception and Status Retry
Handle both network exceptions and HTTP status codes.
Detailed worked examples and patterns
Detailed sections (starting with Advanced Patterns ) live in references/details.md . Read that file when the navigation summary above is insufficient.
Best Practices Summary
1. Retry only transient errors Don't retry bugs or authentication failures
2. Use exponential backoff Give services time to recover
3. Add jitter Prevent thundering herd from synchronized retries
4. Cap total duration stop after attempt(5) stop after delay(60)
5. Log every retry Silent retries hide systemic problems
6. Use decorators Keep retry logic separate from business logic
7. Inject dependencies Make infrastructure testable
8. Set timeouts everywhere Every network call needs a timeout
9. Fail gracefully Return cached/default values for non critical paths
10. Monitor retry rates High retry rates indicate underlying issues