Retries, Timeouts, Backoff, and Circuit Breakers
Set API deadlines, bounded retries, exponential backoff, and circuit breakers so transient failures recover without multiplying load or hiding outages.
API clients can recover from brief faults with bounded retries, per-attempt timeouts, jittered backoff, and a circuit breaker for failures that persist. The examples show how to classify retryable responses, keep attempts inside one caller deadline, protect writes with idempotency, and close retry response bodies. Use the decision points, failure scenarios, and observability checklist to set limits that give a dependency room to recover without turning a slowdown into a traffic spike.
Retries, Timeouts, Backoff, and Circuit Breakers
Introduction
Suppose a client makes three attempts and gives each one a fresh 500 ms timeout. A slow dependency can keep the caller waiting for 1.5 seconds, plus backoff, while concurrent retries add more load to the same service. An overall deadline keeps every attempt and wait inside one request budget.
// Per-attempt timeouts can exceed the caller's total budget.
for (let attempt = 0; attempt < 3; attempt++) {
await fetch(url, { signal: AbortSignal.timeout(500) });
}
This guide shows how to classify retryable failures, cap attempts, use jittered backoff, and combine those controls with a circuit breaker. The goal is to recover from brief faults while failing quickly when the dependency stays unhealthy.
Backoff and breaker behavior
Exponential backoff increases the delay after each failed attempt, often with jitter to avoid many clients retrying together. A practical policy caps both the number of attempts and maximum delay. Honor a valid Retry-After response when it fits within the caller’s deadline. Never use unbounded retries in a request handler: they consume threads, connections, and the user’s time.
A circuit breaker counts failures over a window. In the open state, it rejects requests quickly instead of sending them to a failing dependency. After a cooldown, it permits a small number of probes in a half-open state. A breaker is not a retry mechanism; it limits pressure and gives the dependency room to recover.
The flow below combines the caller’s deadline and retry budget with the breaker lifecycle. Every retry returns through the breaker check, so a breaker that opens during an attempt can stop the next call.
flowchart TD
A[Incoming call] --> B[Set overall deadline and retry budget]
B --> C{Breaker state?}
C -->|Closed| E[Check remaining deadline]
C -->|Open| N[Return fast failure]
C -->|Half-open| F{Probe slot available?}
F -->|Yes| E
F -->|No| N
E --> T{Time remains?}
T -->|No| N
T -->|Yes| G[Call dependency with timeout capped by remaining budget]
G --> H{Attempt succeeds?}
H -->|Yes| I[Record success; close breaker after successful probe]
I --> J[Return response]
H -->|No| K[Classify failure and record breaker result]
K --> L{Threshold reached or probe failed?}
L -->|Yes| M[Open breaker until cooldown]
M --> N[Return terminal failure]
M -->|Cooldown expires| F
L -->|No| O{Retryable and attempts remain?}
O -->|No| N
O -->|Yes| P[Choose capped exponential backoff with jitter]
P --> Q{Wait fits deadline and retry budget?}
Q -->|No| N
Q -->|Yes| R[Consume retry budget and wait]
R --> S{Breaker still closed?}
S -->|Yes| E
S -->|No| C
Trade-Off Table
| Policy choice | Operational effect | Latency and load trade-off |
|---|---|---|
| Fixed backoff | Waits the same interval after each retry; easy to predict and tune for a brief, known interruption. | Keeps recovery timing simple, but clients failing together can retry together and create a traffic spike. |
| Exponential backoff with jitter | Grows the wait after each failure and randomizes it; cap the delay and attempt count. | Spreads retry traffic and reduces pressure, while later attempts can add more tail latency. |
| Shared retry budget | Limits total retries across requests or over a time window, alongside the per-request attempt cap. | Prevents retries from consuming most of the dependency capacity, but some requests fail sooner when the budget runs out. |
| Per-call timeout plus circuit breaker | The timeout bounds one attempt; the breaker rejects calls quickly after sustained failures and allows limited probes after cooldown. | Short timeouts and sensitive breaker thresholds shed load sooner, but can cut off slow or recovering calls; loose limits increase waiting and pressure. |
Implementation snippet
This simplified function retries only explicitly retryable responses and respects an overall budget. Production code should also classify transport errors and account for cancellation.
async function fetchWithRetry(
url: string,
budgetMs: number,
): Promise<Response> {
const started = Date.now();
for (let attempt = 0; attempt < 3; attempt++) {
const remaining = budgetMs - (Date.now() - started);
if (remaining <= 0) throw new Error("request deadline exceeded");
const response = await fetch(url, {
signal: AbortSignal.timeout(remaining),
});
const remainingAfterResponse = budgetMs - (Date.now() - started);
if (remainingAfterResponse <= 0)
throw new Error("request deadline exceeded");
if (![502, 503, 504].includes(response.status) || attempt === 2)
return response;
await response.body?.cancel();
const delay = Math.min(
100 * 2 ** attempt + Math.random() * 100,
remainingAfterResponse,
);
await new Promise((resolve) => setTimeout(resolve, delay));
}
throw new Error("unreachable");
}
When to use and when not to
Use retries for transient errors where a later attempt may succeed and the operation is safe to repeat. Use circuit breakers when a dependency can remain unhealthy long enough that continued calls waste resources. Do not retry permanent client errors, do not retry non-idempotent writes blindly, and do not set a timeout without checking the service’s normal latency and the user’s full request budget.
Production failure scenarios and mitigations
A fleet restarts and all clients retry at once. Add jitter and server-side rate limits. A request times out after the server commits a payment, then the client retries and charges twice. Require idempotency keys and persist their result. A breaker opens due to a few slow but successful calls. Separate latency and error thresholds, inspect per-operation metrics, and test the breaker under load before rollout. Nested retries at several layers can create dozens of attempts; assign retry ownership to one layer and propagate a shared deadline.
Observability checklist
- Measure attempt count, final outcome, and time spent waiting between attempts.
- Track timeout rates and latency percentiles by dependency and operation.
- Export breaker state transitions and rejected-call counts.
- Alert on sustained retry volume, not only final request failures.
- Include trace and correlation IDs while avoiding secrets in logs.
Security notes and pitfalls
Retries can amplify a denial-of-service condition, so apply concurrency limits and bounded queues. Treat remote error bodies as untrusted input, cap their size, and avoid copying sensitive payloads into logs. Do not interpret a client-side timeout as a cancellation guarantee. Keep breaker state local to the dependency boundary; a single global breaker can let one failing route block unrelated work.
Quick Recap Checklist
- Set a request deadline and fit all attempts within it.
- Retry only transient failures, with an attempt cap and randomized backoff.
- Make writes idempotent or otherwise safe to repeat.
- Use a circuit breaker to limit calls during sustained dependency failures.
Interview Questions
Further Reading
- Synchronous Calls, Asynchronous Messaging, and Webhooks — choose communication patterns and acknowledgement behavior.
- Idempotency, Deduplication, and Safe Replays — protect writes from duplicate effects.
- Rate Limits, Quotas, and Consumer Fairness — control request volume at API boundaries.
- Partial Failure, Ordering, and Eventual Consistency — reason about distributed failures and recovery.
- AWS Builders’ Library: Timeouts, retries, and backoff with jitter — practical guidance on selecting and coordinating retry controls.
- Google SRE: Addressing Cascading Failures — how overloaded dependencies and retries can amplify an outage.
- RFC 9110: HTTP Semantics — HTTP status and retry-related response semantics.
Conclusion
Treat retries as a scarce recovery tool. Bound them with time and attempt limits, add jitter, protect writes from duplication, and use a circuit breaker when failures persist. The goal is to recover from brief faults without turning them into a traffic storm.
Category
Related Posts
Idempotency, Deduplication, and Safe Replays
Design idempotent API operations and deduplication records so clients can retry after timeouts without creating duplicate payments, jobs, or updates.
Network Latency, Timeouts, and Failure
Learn how latency, bandwidth, and jitter shape backend requests, then set useful timeouts, bounded retries, and failure handling without amplifying outages.
API Clients, Servers, and Network Boundaries Explained
Understand what API clients and servers each own, how network boundaries fail, and how timeouts, retries, and trust boundaries shape reliable integrations.