Network Latency, Timeouts, and Failure
Learn how latency, bandwidth, and jitter shape backend requests, then set useful timeouts, bounded retries, and failure handling without amplifying outages.
Network calls can outlast a caller or finish after a response is lost, so backend operations need deadlines and per-dependency budgets. This guide explains where bounded retries, jitter, idempotency keys, circuit breakers, and asynchronous queues help, while covering risks such as retry amplification and SSRF from caller-supplied URLs. Use the examples and observability checks to spot tail latency, contain excess load, and handle uncertain outcomes safely.
Network Latency, Timeouts, and Failure
Introduction
A request may cross several services before it returns. Any hop can slow down or disappear, even if the last thousand calls worked. Treat every network call as fallible.
That does not mean every error deserves a retry. First understand what kind of delay or failure occurred, how much time the caller can still wait, and whether repeating the operation is safe.
This guide explains deadlines, bounded retries, idempotency, and circuit breakers, then shows how to diagnose common network failures without adding load to an unhealthy dependency.
What a timeout means
A timeout limits how long the caller waits; it does not prove the server stopped. The server may commit a write whose response was lost. Retrying a payment without an idempotency key can charge twice.
Set timeouts from the caller’s deadline. An API with 1.5 seconds to answer needs shorter downstream budgets, leaving time to validate and respond. Longer timeouts tie up sockets after the caller has left.
When to use retries, and when not to
Retry transient failures, such as connection resets or HTTP 503, only for idempotent work or writes protected by an idempotency key. Use exponential backoff and random jitter.
Do not retry validation or authentication errors, most 4xx responses, or unsafe writes with unknown outcomes. Skip retries when the deadline is nearly spent. A circuit breaker can stop calls during sustained failures.
Use a circuit breaker when repeated dependency failures are wasting caller time or adding load; a single slow response is usually a reason to enforce the deadline, not to open the breaker. Put work on an asynchronous queue when the caller can receive an accepted or pending response and the result can arrive later, such as report generation or email delivery. Keep the queue bounded, make consumers safe to retry, and avoid this pattern when the caller needs the result before it can continue.
Request and dependency flow
sequenceDiagram
participant Client
participant API as Backend API
participant Pay as Payment service
Client->>API: Create order (deadline 1500 ms)
API->>Pay: Charge (budget 700 ms)
Pay-->>API: Timeout or response
Note over API,Pay: Retry only if the outcome is safe to repeat
API-->>Client: Result or bounded failure
A bounded request example
This idempotent GET gets at most three attempts and a two-second budget. The code retries selected status codes and likely network or timeout errors, not programming errors.
const sleep = (ms: number): Promise<void> =>
new Promise((resolve) => setTimeout(resolve, ms));
async function getWithRetries(url: string): Promise<Response> {
const parsed = new URL(url);
if (!["http:", "https:"].includes(parsed.protocol)) {
throw new Error("Unsupported URL protocol");
}
const deadline = Date.now() + 2_000;
const maxAttempts = 3;
for (let attempt = 1; attempt <= maxAttempts; attempt += 1) {
const remainingMs = deadline - Date.now();
if (remainingMs <= 0) throw new Error("Request deadline exceeded");
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), remainingMs);
try {
const response = await fetch(url, { signal: controller.signal });
const retryable = [429, 502, 503, 504].includes(response.status);
if (!retryable || attempt === maxAttempts) return response;
await response.body?.cancel();
} catch (error) {
const transient =
error instanceof TypeError ||
(error instanceof DOMException && error.name === "AbortError");
if (!transient || attempt === maxAttempts) throw error;
} finally {
clearTimeout(timer);
}
const backoffMs = 100 * 2 ** (attempt - 1) + Math.random() * 100;
const pauseMs = Math.min(backoffMs, deadline - Date.now());
if (pauseMs <= 0) throw new Error("Request deadline exceeded");
await sleep(pauseMs);
}
throw new Error("Request attempts exhausted");
}
The scheme check does not prevent server-side request forgery (SSRF). If a caller can supply the URL, enforce an explicit destination allowlist as well. Aborting a fetch stops the client from waiting; the server may still finish. For writes, use an idempotency key and check the outcome before repeating work. Never put credentials or payloads in retry logs.
Failure modes and trade-offs
| Choice | Benefit | Cost or risk |
|---|---|---|
| Short timeout | Releases request resources quickly | Can reject healthy but slow work |
| Longer timeout | Allows slow dependencies to finish | Holds sockets and workers, increasing queueing |
| Bounded retry with jitter | Recovers from brief faults | Adds latency and extra dependency load |
| No retry | Keeps load predictable | Exposes callers to transient faults |
| Circuit breaker | Fails fast during sustained trouble | Needs thresholds and a fallback policy |
Failures often combine. Packet loss raises timeouts; immediate retries then add load to a service already falling behind. A dependency may accept a write but lose the response, leaving the caller unsure of the result. Use deadlines, small retry budgets, idempotency keys, and queue limits. Load balancing can route around an unhealthy instance, but cannot fix a dependency that is slow everywhere.
Production Failure Scenarios
Tail-latency overload
Failure: A dependency’s p99 latency rises while its median stays normal. Requests queue behind slow calls until worker pools and sockets fill, pushing more requests past their deadlines.
Detection: Track p95/p99 by dependency alongside in-flight requests, queue depth, and timeout rate. Compare client wait time with server completion time to find work that continues after callers give up.
Mitigation: Set per-hop deadlines, cap concurrent calls, and shed work when capacity is exhausted. Use a fallback only when it returns safe, useful data.
Retries amplify an outage
Failure: Clients retry timeouts and 503 responses immediately. The already saturated dependency receives more requests, slows further, and triggers retries from additional services.
Detection: Watch attempts per original request, retry volume, dependency latency, and circuit-breaker state. If attempts per request rise, retries are adding load to the outage.
Mitigation: Retry at one layer, cap attempts and total retry time, and apply exponential backoff with jitter. Use a retry budget and circuit breaker; protect writes with idempotency keys.
Partial network failure leaves an unknown result
Failure: A backend commits a write, but a connection reset or lost response prevents the caller from seeing success. Other services remain reachable, so the failure may look like an isolated client timeout.
Detection: Compare client timeout/reset traces with server-side completion records and operation status. Alert on a rise in completed operations whose responses did not reach callers.
Mitigation: Give the operation an idempotency key and let callers check its status before retrying. Never repeat an unsafe write just because the response was lost.
Observability checklist
- Record dependency, operation, attempt, and outcome without secrets or payloads.
- Measure tail latency, timeout and retry rates, and error class by dependency.
- Propagate traces; alert when slow spans or retry volume consume request budgets.
- Compare client timeouts with server completion and cancellation metrics.
Security and Compliance Notes
Use TLS and validate certificates. Limit connections and response sizes, especially for user-controlled URLs, to reduce resource exhaustion and server-side request forgery risk. Keep authorization headers and sensitive data out of logs.
For regulated workloads, check the applicable requirements for encryption in transit, log access, audit records, and retention. TLS protects a connection but does not by itself make the service compliant.
Common Pitfalls and Anti-Patterns
Do not retry every exception or non-idempotent write, reuse the same large timeout everywhere, watch only averages, or assume a client abort cancels server work. Retries are part of load management, not a generic error handler.
Quick recap checklist
- Distinguish latency, bandwidth, and jitter; inspect tail latency.
- Give each operation a deadline and each dependency a smaller budget.
- Retry only transient failures, within a small attempt and time limit.
- Use idempotency for writes and jitter to spread retries.
- Observe timeout and retry rates alongside successful latency.
Interview Questions
Expected answer points:
- Propagate the caller's deadline or remaining budget to each dependency instead of giving every hop the full original timeout.
- Reserve time for application work, response validation, and sending the result back to the caller.
- For sequential calls, ensure their budgets and processing time fit within the total deadline; parallel calls still need limits and cancellation behavior.
Expected answer points:
- No. The client can stop waiting while the server continues processing.
- Cancellation must be propagated through the application and dependencies, and each component must honor it.
- For a write with an unknown outcome, check its status or use an idempotency key before retrying.
Expected answer points:
- Use a queue when the caller can receive an accepted or pending result and the work can finish later.
- Keep the queue bounded and make consumers safe to retry so a dependency outage does not create unbounded backlog or duplicate effects.
- It is a poor fit when the caller needs the completed result before it can continue.
Expected answer points:
- Without jitter, many clients that fail together may retry at the same intervals.
- Random delay spreads attempts over time and reduces synchronized load spikes at the recovering dependency.
- Backoff still needs an attempt limit and the caller's overall deadline.
Further Reading
- HTTP and HTTPS Protocol — connection setup, transport versions, and request behavior.
- Circuit Breaker Pattern — stop repeated calls to a dependency that is failing.
- Load Balancing — route around unhealthy backends and manage connection health.
- Google SRE: Addressing Cascading Failures — guidance on retry backoff and overload.
Conclusion
Network calls can stall, fail halfway through, or succeed without delivering a response. Give them deadlines, keep retries bounded, and make repeated writes safe. Those small rules keep one slow dependency from consuming the time and capacity of the rest of a backend.
Category
Related Posts
Network Observability: Signals for Reliable Services
Track network health across hosts, DNS, paths, proxies, and requests. Learn which signals help diagnose failures without confusing telemetry with service SLOs.
Background Jobs, Scheduling, and Worker Pools
Design background jobs and worker pools with bounded concurrency, safe retries, scheduling, and production checks that keep slow work out of request paths.
Debugging Backend Applications
Use a repeatable backend debugging workflow to reproduce failures, inspect evidence, test one hypothesis at a time, and verify fixes safely in production.