Python resilience patterns including automatic retries, exponential backoff, timeouts, and fault-tolerant decorators. Use when adding retry logic, implementing timeouts, building fault-tolerant services, or handling transient failures.
Permissions
Just instructions. No commands, network, or file access.
Files
SKILL.md
python-resilience
Python resilience patterns including automatic retries, exponential backoff, timeouts, and fault-tolerant decorators. Use when adding retry logic, implementing timeouts, building fault-tolerant services, or handling transient failures.
Python Resilience Patterns
Build fault-tolerant Python applications that gracefully handle transient failures, network issues, and service outages. Resilience patterns keep systems running when dependencies are unreliable.
When to Use This Skill
Adding retry logic to external service calls
Implementing timeouts for network operations
Building fault-tolerant microservices
Handling rate limiting and backpressure
Creating infrastructure decorators
Designing circuit breakers
Core Concepts
1. Transient vs Permanent Failures
Retry transient errors (network timeouts, temporary service issues). Don't retry permanent errors (invalid credentials, bad requests).
2. Exponential Backoff
Increase wait time between retries to avoid overwhelming recovering services.
3. Jitter
Add randomness to backoff to prevent thundering herd when many clients retry simultaneously.
4. Bounded Retries
Cap both attempt count and total duration to prevent infinite retry loops.
Use the tenacity library for production-grade retry logic. For simpler cases, consider built-in retry functionality or a lightweight custom implementation.
Detailed sections (starting with ## Advanced Patterns) live in references/details.md. Read that file when the navigation summary above is insufficient.
Best Practices Summary
Retry only transient errors - Don't retry bugs or authentication failures
Use exponential backoff - Give services time to recover
Add jitter - Prevent thundering herd from synchronized retries
Cap total duration - stop_after_attempt(5) | stop_after_delay(60)
Log every retry - Silent retries hide systemic problems
Use decorators - Keep retry logic separate from business logic
Inject dependencies - Make infrastructure testable
Set timeouts everywhere - Every network call needs a timeout
Fail gracefully - Return cached/default values for non-critical paths
Monitor retry rates - High retry rates indicate underlying issues