python-resilience

Python resilience patterns including automatic retries, exponential backoff, timeouts, and fault-tolerant decorators. Use when adding retry logic, implementing timeouts, building fault-tolerant services, or handling transient failures.

By wshobson · 10,197 installs

npx skills add wshobson/agents --skill python-resilience

Source repository · Upstream listing

Python Resilience Patterns Build fault tolerant Python applications that gracefully handle transient failures, network issues, and service outages. Resilience patterns keep systems running when dependencies are unreliable. When to Use This Skill Adding retry logic to external service calls Implementing timeouts for network operations Building fault tolerant microservices Handling rate limiting and backpressure Creating infrastructure decorators Designing circuit breakers Core Concepts 1. Transient vs Permanent Failures Retry transient errors (network timeouts, temporary service issues). Don't retry permanent errors (invalid credentials, bad requests). 2. Exponential Backoff Increase wait time between retries to avoid overwhelming recovering services. 3. Jitter Add randomness to backoff to prevent thundering herd when many clients retry simultaneously. 4. Bounded Retries Cap both attempt count and total duration to prevent infinite retry loops. Quick Start Fundamental Patterns Pattern 1: Basic Retry with Tenacity Use the tenacity library for production grade retry logic. For simpler cases, consider built in retry functionality or a lightweight custom implementation. Pattern 2: Retry Only Appropriate Errors Whitelist specific transient exceptions. Never retry: ValueError , TypeError These are bugs, not transient issues AuthenticationError Invalid credentials won't become valid HTTP 4xx errors (except 429) Client errors are permanent Pattern 3: HTTP Status Code Retries Retry specific HTTP status codes that indicate transient issues. Pattern 4: Combined Exception and Status Retry Handle both network exceptions and HTTP status codes. Detailed worked examples and patterns Detailed sections (starting with Advanced Patterns ) live in references/details.md . Read that file when the navigation summary above is insufficient. Best Practices Summary 1. Retry only transient errors Don't retry bugs or authentication failures 2. Use exponential backoff Give services time to recover 3. Add jitter Prevent thundering herd from synchronized retries 4. Cap total duration stop after attempt(5) stop after delay(60) 5. Log every retry Silent retries hide systemic problems 6. Use decorators Keep retry logic separate from business logic 7. Inject dependencies Make infrastructure testable 8. Set timeouts everywhere Every network call needs a timeout 9. Fail gracefully Return cached/default values for non critical paths 10. Monitor retry rates High retry rates indicate underlying issues