API Gateway: Single Entry Point for Microservices
Learn how API gateways work, when to use them, architecture patterns, failure scenarios, and implementation strategies for production microservices.
An API gateway gives clients one entry point for backend services and can apply shared edge policies such as credential checks, routing, and rate limits. This guide compares gateways with BFFs and service meshes, explains failure handling and measured capacity planning, and walks through a Node.js implementation. It also clarifies that each service remains responsible for authenticating trusted callers and authorizing access to its own resources.
API Gateway: The Single Entry Point for Microservices Architecture
Introduction
A mobile product screen needs catalog details, live inventory, and a price. Without a gateway, the app must discover and call three services, handle each service’s authentication and errors, and combine the responses. With a gateway, it can make one request such as GET /v1/products/42; the gateway checks the caller’s credential, applies edge policies, and routes or aggregates the work. The catalog and inventory services still decide whether that caller may read the requested data.
This guide covers gateway responsibilities, rate limiting, authentication, BFFs, failure handling, capacity planning, and a Node.js implementation. It also shows where gateway policy ends and service-owned authorization begins.
Core Concepts
Request Flow
The sequence diagram shows a request traveling through an API gateway from the moment a client sends it to the moment the gateway returns a response. It covers the two most common scenarios: a cache hit that short-circuits the backend call, and a cache miss that requires the gateway to fan out to multiple services, aggregate their responses, and return a consolidated result.
The first scenario is a product lookup. The gateway validates the JWT and checks rate limits before doing anything else, then hits Redis. For publicly readable product data, a cache hit can skip the product service. For user- or tenant-scoped data, authorize before returning a cached entry and include the authorized scope in the cache key. The second scenario is order creation, which requires the gateway to check product availability before creating the order, then wait for the order service to confirm. The client makes one call; the gateway makes three.
Both paths start the same way. Same endpoint, same security checks. They diverge only after the cache lookup. This is the real benefit of putting a gateway in front of everything: clients deal with one address, one auth mechanism, and one error format. The complexity of backend service discovery, protocol translation, and response aggregation stays hidden.
sequenceDiagram
participant Client
participant Gateway as API Gateway
participant Auth as Auth Service
participant Catalog as Product Service
participant Order as Order Service
participant Cache as Redis Cache
Client->>Gateway: POST /api/products/123
Gateway->>Gateway: Extract JWT, Rate Limit Check
Gateway->>Auth: Validate Token
Auth-->>Gateway: Token Valid
Gateway->>Cache: Check Cache
Cache-->>Gateway: Cache Hit
Gateway-->>Client: Product JSON
Client->>Gateway: POST /api/orders
Gateway->>Gateway: Extract JWT, Rate Limit Check
Gateway->>Auth: Validate Token
Auth-->>Gateway: Token Valid
Gateway->>Catalog: Check Product Availability
Catalog-->>Gateway: Available
Gateway->>Order: Create Order
Order-->>Gateway: Order Created
Gateway-->>Client: Order Confirmation
Gateway Internal Components
The gateway processes every inbound request through a layered pipeline. The diagram below breaks this into two logical groups: the security layer that every request passes through before routing decisions are made, and the routing layer that determines where the request goes and how the response comes back.
The gateway can handle shared edge concerns such as TLS termination, credential validation, and broad rate limits before traffic reaches backend services. Services must still authorize access to the specific resources they own; a gateway role check cannot replace a tenant or object-level permission check.
The routing layer takes over after security checks pass. It looks up backend instances in the service registry, applies routing rules (path rewriting, header injection, canary weight splitting), forwards the request, and handles the response. If the gateway is aggregating responses from multiple services, it calls them in parallel, waits, merges the results, and sends a single reply to the client. Metrics get emitted at every step.
graph TD
A[Client Request] --> B[TLS Termination]
B --> C[Authentication]
C --> D[Authorization]
D --> E[Rate Limiting]
E --> F[Request Routing]
F --> G[Service Discovery]
G --> H[Backend Service]
H --> F
F --> I[Response Aggregation]
I --> J[Metrics Collection]
J --> K[Client Response]
subgraph Security Layer
C
D
E
end
subgraph Routing Layer
F
G
end
Route Policy Example
A route should make its upstream, identity requirement, and rate-limit key explicit. Match specific paths before broad catch-all routes, and strip or rewrite only the prefix the upstream expects. A gateway can enforce coarse access policy, but the service still checks ownership of the requested resource.
| Method and client path | Upstream path | Gateway policy | Rate-limit key |
|---|---|---|---|
GET /v1/products/{id} |
GET /products/{id} on catalog |
Require a valid access token; catalog checks product visibility | consumer_id |
POST /v1/orders |
POST /orders on orders |
Require a valid token; orders checks role, tenant, and purchase permissions | consumer_id |
GET /health |
Gateway health handler | Public, with no backend route | None |
For example, the gateway can strip /v1 before forwarding GET /v1/products/42 as GET /products/42. It should reject an unknown method or path rather than send it to a permissive default upstream. If the gateway uses client IP for anonymous limits, trust X-Forwarded-For only when a known proxy has set it; otherwise clients can choose the value themselves.
Failure Flow
The gateway sits in the critical path for every request, which means how it handles failures determines whether downstream services ever see traffic at all. The failure decision tree maps every possible error condition to the most appropriate HTTP status code and tells the client what to do next.
The first check is whether the gateway itself is operational. If the gateway is unavailable (instance crash, OOM, deployment in progress), clients get a 503. This is terminal; the client can only retry after a backoff. If the gateway is up, the request proceeds to authentication.
Auth failures return 401. Rate limit violations return 429 with a Retry-After header. Backend service unavailability returns 502. Service timeouts return 504. Each maps to a specific failure mode and tells the client something useful about what happened and when to retry.
The design principle is fail-fast. Return an error as soon as you know the request cannot succeed, not after running through the entire pipeline only to fail at the end. This protects gateway resources and gives clients actionable signals.
graph TD
A[Client Request] --> B{Gateway Available?}
B -->|No| C[Return 503 Service Unavailable]
B -->|Yes| D{Auth Passed?}
D -->|No| E[Return 401 Unauthorized]
D -->|Yes| F{Rate Limit OK?}
F -->|No| G[Return 429 Too Many Requests]
F -->|Yes| H{Backend Service Available?}
H -->|No| I[Return 502 Bad Gateway]
H -->|Yes| J{Request Valid?}
J -->|No| K[Return 400 Bad Request]
J -->|Yes| L[Forward to Service]
L --> M{Service Timeout?}
M -->|Yes| N[Return 504 Gateway Timeout]
M -->|No| O[Return Service Response]
When to Use / When Not to Use
| Scenario | Recommendation |
|---|---|
| Multiple backend services need unified access control | Use API Gateway |
| Mobile, web, and third-party clients consume the same APIs | Use API Gateway |
| You need centralized rate limiting and throttling | Use API Gateway |
| Service aggregation is required for client convenience | Use API Gateway |
| Single monolithic application with no external clients | Do NOT use API Gateway |
| Services are tightly coupled and share a deployment unit | Do NOT use API Gateway |
| Ultra-low latency is critical (gateway adds ~1-3ms) | Consider alternatives |
| Simple CRUD application with one or two services | Consider direct service calls |
When TO Use an API Gateway
- Unified client access: Your mobile app, web app, and third-party integrations all hit different services. Without a gateway, clients need to know about every service endpoint, certificate, and authentication mechanism.
- Shared authentication and authorization: You want a single place to validate JWTs, check permissions, and reject unauthorized requests before they reach your services.
- Rate limiting at the edge: You need to protect your services from traffic spikes, abusive clients, or accidental misconfiguration without adding this logic to every service.
- Protocol translation: Your mobile clients use REST, but your internal services might use gRPC or WebSocket. The gateway translates between them.
- Request aggregation: A mobile screen needs data from three different services. Without aggregation in the gateway, the client makes three separate calls with associated latency and complexity.
When NOT to Use an API Gateway
- Adding unnecessary hops: If your system is a simple monolith or a handful of tightly coordinated services, the gateway introduces latency without meaningful benefit.
- Bypassing for internal services: In some architectures, internal services behind the gateway still need to call each other directly. The gateway becomes a bottleneck rather than a helper.
- Single-purpose applications: A data processing pipeline with no external clients does not need a gateway.
- Latency-sensitive paths: Every request going through the gateway adds 1-3ms. For extremely latency-sensitive use cases, this matters.
Rate Limiting Algorithms
Not all rate limiting works the same way. The algorithm you pick affects burst tolerance, memory usage, and how fairly limits get enforced across clients.
| Algorithm | How it works | Burst tolerance | Memory | Best for |
|---|---|---|---|---|
| Fixed Window | Count requests per fixed time window (e.g., 100/min) | High at window boundary | Low | Simple cases, approximate enforcement |
| Sliding Window Log | Store timestamps per request, count within rolling window | Accurate, no boundary burst | High | Exact enforcement, lower QPS APIs |
| Sliding Window Counter | Weighted interpolation between adjacent windows | Low | Low | Balance of accuracy and memory |
| Token Bucket | Tokens added at fixed rate; each request consumes one | Controlled bursts allowed | Low | APIs with bursty clients |
| Leaky Bucket | Requests queue up and process at fixed rate | No burst — queue or drop | Low | Smoothing traffic to backends |
The fixed window edge case
Fixed windows have a known problem: clients can effectively double their rate by sending requests at the end of one window and the start of the next.
Window 1 (0-60s): 90 requests at t=59s
Window 2 (60-120s): 90 requests at t=61s
Effective rate: 180 requests in 2 seconds, both windows satisfied
Sliding window approaches fix this. The log variant is exact but stores one entry per request. The counter variant approximates the sliding window using weights between two adjacent fixed windows — much cheaper on memory with acceptable accuracy.
Token bucket in practice
Token bucket is the most common choice for API gateways. It allows short bursts up to bucket capacity while enforcing a long-term average rate. Here is a Redis implementation using atomic operations:
async function tokenBucketAllow(userId, maxTokens, refillRate) {
const key = `rate:${userId}`;
const now = Date.now();
const bucket = await redis.hgetall(key);
const tokens = bucket ? parseFloat(bucket.tokens) : maxTokens;
const lastRefill = bucket ? parseFloat(bucket.lastRefill) : now;
// Refill tokens based on elapsed time
const elapsed = (now - lastRefill) / 1000;
const refilled = Math.min(maxTokens, tokens + elapsed * refillRate);
if (refilled < 1) {
const retryAfter = Math.ceil((1 - refilled) / refillRate);
return { allowed: false, retryAfter };
}
await redis.hset(key, { tokens: refilled - 1, lastRefill: now });
await redis.expire(key, 3600);
return { allowed: true, remaining: Math.floor(refilled - 1) };
}
This sketch shows the bucket math, not a safe concurrent Redis implementation: HGETALL, HSET, and EXPIRE are separate operations, so two gateway instances can both read the same token count and allow more requests than intended. In production, perform read, refill, decision, update, and expiry in one Redis Lua script or another atomic operation. Shared Redis state makes limits consistent across instances only when each update is atomic; local in-memory limits remain per-process.
When Redis is unavailable, choose behavior by endpoint risk. Failing open keeps ordinary reads available but removes the quota; failing closed preserves a spend or abuse limit but can reject legitimate requests. Use a bounded local fallback only if the resulting per-instance limit is acceptable, and return 429 with Retry-After when a request is over quota. Clients should honor that delay and add jitter to retries so a limit response does not trigger a synchronized retry surge.
Authentication Strategies at the Gateway
The gateway can validate credentials at the edge so invalid requests are rejected before reaching backend services. Services must still authenticate trusted callers where their boundary requires it and authorize access to the resources they own. Each gateway strategy has a different trade-off between revocation speed, overhead, and operational complexity.
| Strategy | How it works | Revocation | Overhead | Best for |
|---|---|---|---|---|
| API Keys | Static key in header or query string | Immediate (delete) | Very low | Machine-to-machine, third-party devs |
| JWT (stateless) | Signed token decoded locally at gateway | Requires blocklist | Very low | Internal services, short-lived tokens |
| OAuth 2.0 + JWT | Token from auth server, decoded or introspected | Via introspection | Medium | User-facing APIs |
| mTLS | Mutual TLS certificates both sides | CRL / OCSP | High | Service-to-service, regulated envs |
| Session tokens | Opaque token looked up in session store per request | Immediate | Medium | Traditional web apps |
The JWT revocation problem
Stateless JWT validation is fast because the gateway decodes the token locally without calling another service. The problem: you cannot revoke a JWT before it expires.
If a user logs out and you issued a JWT with a one-hour TTL, that token stays valid for up to an hour.
Two practical mitigations:
- Short TTL plus refresh tokens: Issue JWTs with 5-15 minute TTLs. Clients use a longer-lived refresh token to get new JWTs. The revocation window equals the TTL.
- Token blocklist in Redis: Store revoked token IDs (JTI claim) in Redis with a TTL matching the original JWT TTL. The gateway checks the blocklist on every request. Costs about 1ms per check.
For most applications, short TTLs with refresh tokens are the right call. Blocklists are worth adding if you need immediate revocation — compliance requirements, suspected credential compromise, or account suspension flows.
Backend for Frontend (BFF) Pattern
A Backend for Frontend (BFF) is a specialized gateway instance tailored to a specific client type. Instead of one generic gateway that mobile, web, and partner clients all share, you build separate gateways per client.
graph TD
A[Mobile App] --> B[Mobile BFF]
C[Web App] --> D[Web BFF]
E[Partner API] --> F[Partner Gateway]
B --> G[Product Service]
B --> H[Order Service]
D --> G
D --> H
D --> I[Recommendation Service]
F --> G
BFFs solve the problem of one gateway trying to serve every client’s needs. Mobile apps typically want smaller payloads, fewer fields, and different aggregation than web apps. Without BFF, you end up with a bloated general-purpose gateway that either handles every possible client requirement or pushes aggregation logic into the clients themselves.
| Approach | Complexity | Flexibility | Team ownership |
|---|---|---|---|
| Single gateway | Low | Limited | Centralized platform team |
| BFF per client type | Medium | High | Per-client teams own their BFF |
| BFF per team | High | Very high | Full autonomy, but risk of duplication |
BFF works well when:
- Different client types have significantly different data requirements
- Teams have clear ownership boundaries (mobile team, web team, partner integrations team)
- You have enough traffic to justify separate deployments
It adds complexity when:
- Teams are small and one group would own multiple BFFs
- Clients have mostly overlapping requirements
- Deployment automation is not already mature
Production Failure Scenarios
| Failure Scenario | Impact | Mitigation |
|---|---|---|
| Gateway instance crash | Available capacity drops; remaining instances may overload | Run multiple gateway instances behind a load balancer and size for an instance loss |
| Backend service timeout | The request exceeds its route deadline | Set an explicit deadline and return a 504 when it expires |
| Auth service unavailable | No requests can be validated | Use valid cached signing keys where policy permits; fail closed with 503 if a credential cannot be verified |
| Rate limiter memory exhaustion | Rate limiting fails open | Use Redis-backed rate limiting; set hard limits on memory per tenant |
| Rate-limit store unavailable | Limits may be bypassed or valid requests rejected | Choose fail-open or fail-closed behavior by route risk; alert on store errors and test the configured fallback |
| Gateway misconfiguration | All traffic routing incorrectly | Use version-controlled config; canary deployments for config changes |
| SSL/TLS certificate expiry | HTTPS requests fail | Automate certificate renewal (Let’s Encrypt); alert 30 days before expiry |
| Service discovery returns stale IPs | Requests go to dead instances | Use short TTL in service registry; health checks remove unhealthy instances |
| Request payload too large | Memory exhaustion on gateway | Set max request size limits; reject oversized payloads early |
Common Pitfalls / Anti-Patterns
Pitfall 1: Gateway as a Monolith Proxy
Teams sometimes build the gateway to contain significant business logic, transforming it into another monolith that mirrors the old system. Instead of routing requests, the gateway accumulates conditional logic, data transformations, and workflow orchestration. Over time, it becomes the most critical—and most fragile—component in the stack. A change to pricing logic should not require a gateway deployment.
The symptom is unmistakable: gateway code grows faster than service code, and deployments start requiring gateway changes for every feature. Teams blame the monolith they just replaced and build a new one with extra hops. Business logic in the gateway means every team needs the gateway team to review and deploy their changes, bottlenecking feature velocity.
The fix is strict layering from the start. The gateway owns the network layer: auth headers, route decisions, protocol translation. Backend services own product logic, pricing rules, and domain workflows. When a product team asks to add a feature to the gateway, the question to ask is whether that logic belongs in the gateway or the service. Most of the time it belongs in the service.
Pitfall 2: No Circuit Breaker on Backend Calls
A gateway without circuit breakers can keep work queued against a failing backend until its own connection or memory limits are reached. For example, if payment calls take 30 seconds instead of 50ms, checkout requests accumulate while they wait; enough in-flight work can exhaust gateway resources and affect unrelated routes.
A circuit breaker stops sending calls after a configured failure threshold and returns an error while the dependency recovers. Use it with bounded timeouts and concurrency limits. A fallback is safe only when it preserves the operation’s meaning; cached product descriptions may be acceptable, but stale payment authorization is not.
// Never do this - no timeout, no circuit breaker
const response = await axios.get(`${BACKEND_URL}/data`);
// Always do this
const circuit = new CircuitBreaker(axios.get, {
timeout: 3000,
errorThresholdPercentage: 50,
});
const response = await circuit.fire(BACKEND_URL);
Pitfall 3: Stale Service Discovery
Service discovery gets stale in ways that are hard to detect without active probing. The registry shows five healthy instances, but one of them is in the middle of a rolling deployment and no longer accepting connections. Another three are scheduled for termination but still appear active between health checks. The gateway routes to dead pods while dashboards report everything is green.
Kubernetes endpointslices update when pods change state, but the gateway’s local cache or even the in-cluster discovery client can lag by seconds to minutes. During a rolling deployment or scale-down event, that lag is enough to route traffic to instances that are terminating. Clients see connection refused while the gateway keeps sending traffic to addresses that no longer have anything listening.
Watch-based discovery handles this better than polling. Instead of asking the registry for a list every N seconds, the registry pushes changes to subscribers immediately. In Kubernetes, use the endpointslices API with watch streams rather than list-then-watch patterns. Configure the gateway to treat endpoint changes as signals, not just as data to cache. When an instance disappears from the list, remove it from routing immediately, not after a TTL expires.
Pitfall 4: Authentication Bypass via Direct Service Access
Backend services that skip the gateway for “internal” access create a hole in your security perimeter that attackers look for first. The pattern is common: a developer sets up direct HTTP access for testing, or an internal tool needs to bypass the gateway for speed, or a microservice-to-microservice call skips auth because it is “internal.” Each shortcut opens a path that traffic analysis or a misconfigured firewall exposes to the public internet.
The risk is not hypothetical. Port scans find open management interfaces on cloud instances within minutes of going live. Unauthenticated endpoints—debug pages, internal health endpoints, actuator endpoints with JMX—get scanned and exploited if they face the internet. Rate limiting that does not exist on direct paths means brute force attacks have no throttle. JWT validation that lives in the gateway never runs for traffic that bypasses it.
Enforce network boundaries: keep backend services on private subnets with no public inbound routes, and allow traffic from the gateway or approved internal service identities. Services should still authenticate callers and authorize access to their own resources. For testing, use a separate harness with the relevant checks enabled instead of exposing an unfiltered path.
Pitfall 5: Rate Limiting Without Global State
With four gateway instances handling traffic, a client making 100 requests per minute can hit each instance with 25 requests and pass all of them through. Local in-memory rate limiting is per-process accounting. It works for a single instance. The moment you scale horizontally, you have a distributed hole in your protection that any client can discover by hitting instances round-robin.
The attack is not sophisticated. A script loops through instance IPs or uses connection pooling with the same key to distribute load across gateway processes. For authenticated endpoints, token buckets tied to user ID or API key leak if the gateway does not share state. For anonymous rate limits, IP-based limits split across instances give each one a fresh counter. The fix is simple: store counters in Redis, not in process memory. Every gateway instance reads from and writes to the same distributed state, so the limit is enforced globally, not per-instance.
Real-world Failure Scenarios
These examples show gateway-related failure patterns and the kinds of operational changes teams used to address them.
| Incident | What Happened | Root Cause | Resolution |
|---|---|---|---|
| Cloudflare API outage (2022) | Edge API endpoints returned 502 errors for ~30 minutes | A misconfigured authentication module in their API gateway layer rejected valid requests after a rule deployment | Rollback of gateway configuration rules; staged deployment process with canary testing introduced |
| AWS API Gateway throttling cascade (2020) | Downstream services saw traffic spikes as clients retried after hitting rate limits | Clients received 429 errors, retried immediately, and amplified traffic 3-5x | Implemented exponential backoff with jitter on retry logic; added client-side rate limit awareness |
| Stripe gateway timeout chain | Payment processing API returned timeouts during peak traffic | Gateway had 30s default timeouts; a slow downstream auth service caused timeout cascades | Reduced gateway timeouts to 5s; implemented circuit breakers with fallback responses |
| GitHub API gateway misroute | Internal services received requests with wrong routing headers | A gateway configuration deployment caused routes to be incorrectly rewritten | Configuration validation pipeline added before deployments; route testing in staging |
| Netflix API gateway split-brain | Some API requests succeeded, others failed during regional failover | Gateway instances were not synchronized during failover, serving stale routing tables | Implemented consistent hashing for route lookups; session affinity during failover |
How Incident Response Changes with an API Gateway
When an incident starts, compare gateway status and latency with upstream service timings to locate the failing hop. Trace a sample request through authentication, rate limiting, routing, and backend calls; then check gateway readiness separately from backend health. If symptoms began after a route or policy change, roll back that change and verify recovery with a known request before applying a fix. Keep request IDs and gateway logs available so the service team can correlate the same failures on its side.
Trade-off Analysis
| Factor | With API Gateway | Without API Gateway |
|---|---|---|
| Latency | +1-3ms per request | Baseline |
| Consistency | Centralized auth/rate limiting | Duplicated per service |
| Cost | Gateway instances + operation | No additional cost |
| Complexity | Centralized logic, single config | Distributed logic, multiple configs |
| Operability | Single point to monitor | Monitor each service separately |
| Client complexity | Low (one endpoint) | High (manage multiple endpoints) |
| Debugging | Single point to trace | Trace across multiple services |
| Single point of failure | Yes, unless highly available | No (but more complex clients) |
| Flexibility | Limited by gateway capabilities | Full flexibility per service |
Gateway vs Service Mesh
| Aspect | API Gateway | Service Mesh |
|---|---|---|
| Layer | L7 (Application) | L4/L7 (Transport + Application) |
| Scope | North-South traffic (client to service) | East-West traffic (service to service) |
| Typical Users | Platform teams, API product teams | DevOps, SRE teams |
| Features | Auth, routing, aggregation, protocol translation | mTLS, retries, circuit breaking |
| Deployment | Sits at edge | Sidecar proxies on each service |
For most architectures, you need both. The API gateway handles external client traffic while a service mesh handles internal service-to-service communication. See Service Mesh for a deep dive.
Capacity Estimation
Assumptions
- Average request size: 2 KB
- Average response size: 16 KB
- Peak QPS: 10,000 requests/second
- Average response time target: 50ms (gateway overhead: 3ms)
Gateway Instance Calculation
Size from a load test of the chosen gateway configuration at the target latency and payload mix. Suppose one instance sustains 2,000 QPS at that latency, and you cap planned load at 70% of that result. Its planning capacity is 1,400 QPS.
For 10,000 peak QPS, eight instances would be needed to carry the load at that cap. To tolerate one instance failing without exceeding the cap, provision nine: the eight remaining instances provide 11,200 QPS of planning capacity. Treat these numbers as an example; TLS, plugins, aggregation, connection limits, and response sizes can change measured throughput. Re-test after material config or workload changes.
Network Bandwidth
Inbound: 10,000 QPS × 2 KB = 20 MB/s = 160 Mbps
Outbound: 10,000 QPS × 16 KB = 160 MB/s = 1.28 Gbps
Total network required: ~1.5 Gbps
Memory (per instance with 2 vCPU)
Connection buffers: 256 MB
Rate limiting state (Redis): Shared across instances
Application heap: 512 MB
Operating system: 256 MB
Total per instance: ~1 GB RAM
Operational Checklists
Core Gateway Practices
- An API gateway provides a single entry point for all client requests, handling auth, routing, rate limiting, and protocol translation.
- Use API gateways when you have multiple services, diverse clients, or need centralized security policy enforcement.
- Avoid gateways when latency is critical, for simple single-service applications, or when the overhead outweighs benefits.
- Always implement circuit breakers, proper timeouts, and health checks when calling backend services.
- Run multiple gateway instances behind a load balancer to avoid single points of failure.
- Log structured data (request ID, latency, status) for debugging; emit metrics for alerting.
- Pick your rate limiting algorithm based on burst tolerance requirements — token bucket works for most cases.
- JWT revocation requires either short TTLs with refresh tokens or a Redis-backed blocklist.
- BFF pattern is worth adding when different client types have significantly different data needs.
Observability Checklist
Metrics to Capture
gateway_requests_total(counter) - Total requests by route, status codegateway_request_duration_seconds(histogram) - Latency by route, percentile bandsgateway_upstream_requests_totalandgateway_upstream_duration_seconds- Upstream status and latency by service and routegateway_active_connections(gauge) - Current concurrent connections- Gateway saturation signals such as worker utilization, queue depth, rejected connections, and open file descriptors
gateway_rate_limit_exceeded_total(counter) - Rate limit violations by clientgateway_backend_errors_total(counter) - Backend service errors by servicegateway_circuit_breaker_state(gauge) - Circuit breaker state by backend- Authentication and authorization outcomes by policy result; count rejects without using credentials, user IDs, or raw tokens as metric labels
Logs to Emit
Each request should emit structured JSON logs:
{
"timestamp": "2026-03-23T10:15:30.123Z",
"requestId": "550e8400-e29b-41d4-a716-446655440000",
"method": "GET",
"route": "/api/products/{productId}",
"statusCode": 200,
"latencyMs": 12,
"upstreamService": "product-service",
"upstreamStatusCode": 200,
"authnOutcome": "accepted",
"authzOutcome": "allowed",
"rateLimitRemaining": 87,
"upstreamLatencyMs": 8
}
Use a request or trace ID to correlate gateway and upstream spans, and generate a new ID when the client does not provide a valid one. Do not blindly trust client-supplied correlation headers. Keep raw paths, query strings, authorization headers, cookies, and request bodies out of routine logs; they can contain identifiers or credentials. If client IP is needed for abuse investigations, restrict access and retention, and apply the site’s privacy policy.
Alerts to Configure
| Alert | Threshold | Severity |
|---|---|---|
| P99 latency > 100ms | 100ms for 5 minutes | Warning |
| P99 latency > 500ms | 500ms for 1 minute | Critical |
| Error rate > 1% | 1% for 5 minutes | Warning |
| Error rate > 5% | 5% for 1 minute | Critical |
| Rate limit violations spike | > 1000/min from single IP | Warning |
| Backend service unavailable | Any backend down > 30s | Critical |
| Certificate expiry < 30 days | Any cert expiring soon | Warning |
Distributed Tracing
The gateway must propagate trace context to backend services:
// Propagate trace headers to backend services
const traceHeaders = {
"X-Request-ID": req.id,
"X-B3-TraceId": req.headers["x-b3-traceid"],
"X-B3-SpanId": req.headers["x-b3-spanid"],
"X-B3-Sampled": req.headers["x-b3-sampled"],
};
await axios.get(`${SERVICE_URL}/products/${id}`, {
headers: { ...traceHeaders, Authorization: req.headers.authorization },
});
Security and Compliance Notes
- TLS 1.2+ termination with modern cipher suites
- JWT validation with proper signature verification
- Rate limiting configured per-client (IP, API key, user ID)
- Request size limits to prevent payload amplification
- Validate header names, count, and total size; reject malformed or unexpected forwarding headers
- Input validation on all request parameters
- Output encoding to prevent XSS in responses
- CORS policy properly configured
- Security headers (HSTS, CSP, X-Frame-Options)
- Audit logging for all authentication/authorization failures
- API key rotation mechanism
- Automate TLS certificate renewal and rotate gateway-to-service credentials and signing keys before expiry; verify old credentials are revoked
- Deprecation notices for older API versions
- Penetration testing performed annually
- DDoS protection at edge (Cloudflare, AWS Shield)
- Backend services unreachable from the public internet; allow the gateway and approved internal service identities
- Test network and DNS paths to confirm clients cannot bypass gateway authentication, authorization, rate limits, or logging
Compliance considerations
- Keep request and response bodies out of routine access logs. Redact credentials and personal data from headers and query strings before exporting logs.
- Set log retention and access rules to match the data policy for each environment; audit access to identity and authorization records.
- Confirm where TLS termination and gateway logs are processed when data residency rules apply. Re-encrypt sensitive traffic from the gateway to backend services.
Implementation Example (Node.js)
Here is a small Express example that demonstrates routing, rate limits, authentication, and circuit breaking:
const crypto = require("node:crypto");
const express = require("express");
const axios = require("axios");
const rateLimit = require("express-rate-limit");
const jwt = require("jsonwebtoken");
const app = express();
// Configuration
const PORT = process.env.PORT || 3000;
const AUTH_SERVICE_URL =
process.env.AUTH_SERVICE_URL || "http://auth-service:8080";
const PRODUCT_SERVICE_URL =
process.env.PRODUCT_SERVICE_URL || "http://product-service:8080";
const ORDER_SERVICE_URL =
process.env.ORDER_SERVICE_URL || "http://order-service:8080";
// Middleware: Parse JSON with size limit
app.use(express.json({ limit: "1mb" }));
// Middleware: Request ID for tracing
app.use((req, res, next) => {
req.id = crypto.randomUUID();
res.setHeader("X-Request-ID", req.id);
next();
});
// Middleware: Rate limiting (Redis-backed in production)
const limiter = rateLimit({
windowMs: 60 * 1000, // 1 minute
max: 100, // 100 requests per minute per IP
message: { error: "Too many requests" },
standardHeaders: true,
legacyHeaders: false,
});
app.use("/api/", limiter);
// Middleware: Authentication
async function authenticate(req, res, next) {
const token = req.headers.authorization?.replace("Bearer ", "");
if (!token) {
return res
.status(401)
.json({ error: "Missing authorization token", requestId: req.id });
}
try {
// In production, use a distributed cache for validation results
const decoded = jwt.verify(token, process.env.JWT_SECRET);
req.user = decoded;
next();
} catch (error) {
return res.status(401).json({ error: "Invalid token", requestId: req.id });
}
}
// Middleware: Authorization
function authorize(...allowedRoles) {
return (req, res, next) => {
if (!req.user || !allowedRoles.includes(req.user.role)) {
return res
.status(403)
.json({ error: "Insufficient permissions", requestId: req.id });
}
next();
};
}
// Health check endpoint
app.get("/health", (req, res) => {
res.json({ status: "healthy", timestamp: new Date().toISOString() });
});
// Route: Product catalog with circuit breaker
const { CircuitBreaker } = require("opossum");
const productCircuit = new CircuitBreaker(
async (productId, requestId) => {
const response = await axios.get(
`${PRODUCT_SERVICE_URL}/products/${productId}`,
{
timeout: 5000,
headers: { "X-Request-ID": requestId },
},
);
return response.data;
},
{
timeout: 5000,
errorThresholdPercentage: 50,
resetTimeout: 30000,
},
);
productCircuit.on("fallback", () => ({
error: "Service temporarily unavailable",
}));
productCircuit.on("timeout", () => ({ error: "Service timeout" }));
app.get("/api/products/:id", authenticate, async (req, res) => {
try {
const product = await productCircuit.fire(req.params.id, req.id);
res.json(product);
} catch (error) {
res.status(502).json({ error: "Bad gateway", requestId: req.id });
}
});
// Route: Create order (aggregates product and order services)
app.post(
"/api/orders",
authenticate,
authorize("user", "admin"),
async (req, res) => {
const { productId, quantity } = req.body;
try {
// Check product availability
const productResponse = await axios.get(
`${PRODUCT_SERVICE_URL}/products/${productId}`,
{ timeout: 3000 },
);
if (!productResponse.data.available) {
return res.status(400).json({ error: "Product not available" });
}
// Create order
const orderResponse = await axios.post(
`${ORDER_SERVICE_URL}/orders`,
{ productId, quantity, userId: req.user.id },
{ timeout: 5000 },
);
res.status(201).json(orderResponse.data);
} catch (error) {
if (error.code === "ECONNABORTED") {
return res.status(504).json({ error: "Gateway timeout" });
}
res.status(502).json({ error: "Failed to create order" });
}
},
);
// Error handling middleware
app.use((err, req, res, next) => {
console.error(`[${req.id}] Unhandled error:`, err);
res.status(500).json({ error: "Internal server error", requestId: req.id });
});
app.listen(PORT, () => {
console.log(`API Gateway listening on port ${PORT}`);
});
Docker Compose for Local Development
version: "3.8"
services:
api-gateway:
build: ./api-gateway
ports:
- "3000:3000"
environment:
- JWT_SECRET=your-secret-key
- AUTH_SERVICE_URL=http://auth-service:8080
- PRODUCT_SERVICE_URL=http://product-service:8080
- ORDER_SERVICE_URL=http://order-service:8080
- REDIS_URL=redis://redis:6379
depends_on:
- redis
redis:
image: redis:7-alpine
ports:
- "6379:6379"
auth-service:
image: your-auth-service-image
ports:
- "8080:8080"
product-service:
image: your-product-service-image
ports:
- "8081:8080"
order-service:
image: your-order-service-image
ports:
- "8082:8080"
Quick Recap Checklist
- Keep business logic in backend services and gateway policy focused on shared edge concerns.
- Choose managed or self-hosted infrastructure based on operational capacity and control needs.
- Apply authentication, rate limits, and timeouts consistently across routes.
- Use separate gateway and service-mesh roles when both external and internal traffic need control.
Interview Questions
An API gateway provides a single entry point for all client requests to backend services. It handles cross-cutting concerns that would otherwise be duplicated across services: authentication, authorization, rate limiting, request routing, protocol translation, and observability.
Without a gateway, clients must know about every service endpoint, manage authentication for each, and handle the complexity of calling multiple services. The gateway simplifies client code and gives you a central place to enforce policies.
In-memory rate limiting lets each gateway instance enforce its own limit independently. A user could make N requests per server. With 10 servers, they effectively get 10N requests. This defeats the purpose of rate limiting.
Redis-backed rate limiting uses shared global state. All gateway instances consult the same Redis counter, ensuring consistent enforcement regardless of which instance handles the request.
Redis also handles atomic operations — INCR and EXPIRE work together to increment and auto-expire counters without race conditions. The latency cost (1-2ms) is acceptable for most gateway use cases.
Gateway instance crash: all traffic fails. Mitigate by running multiple instances behind a load balancer with health checks. Gateway instances should be stateless — store session state in Redis, not local memory.
Backend service timeout: threads pile up waiting. Mitigate with aggressive timeouts (5 seconds or less) and circuit breakers. Backend service unavailable: return 502 Bad Gateway immediately rather than waiting. If an auth dependency is unavailable, use valid cached signing keys where policy permits and fail closed with 503 when a credential cannot be verified.
SSL/TLS certificate expiry: all HTTPS requests fail. Automate certificate renewal with Let's Encrypt or similar. Alert 30 days before expiry.
Typical overhead is 1-3ms per request. For most web applications, where backend services respond in tens to hundreds of milliseconds, this is negligible. The gateway's TLS termination, authentication checks, and routing add up to a small fraction of total latency.
For ultra-low-latency applications (high-frequency trading, real-time gaming), 1-3ms matters. In these cases, consider whether a gateway is necessary or if clients can call services directly with appropriate SDKs.
The gateway can actually reduce latency in some cases: response caching eliminates backend calls, and connection pooling to backends amortizes connection setup costs.
TLS termination at the gateway with modern cipher suites only. JWT validation with signature verification before forwarding requests. Backend services should be unreachable from the public internet and accept traffic only from the gateway or approved internal service identities. Each service should still enforce authorization for its own resources.
DDoS protection at the edge — Cloudflare, AWS Shield, or similar. Rate limiting prevents abuse. Request size limits prevent payload amplification attacks. Input validation prevents injection attacks. Audit logging for authentication and authorization failures for compliance.
The gateway can translate between client-facing protocols (REST, GraphQL) and internal protocols (gRPC, WebSocket). It handles content type negotiation via Accept headers, translating between XML and JSON if needed.
Protocol translation lives at the gateway, not in backend services. Backend services speak their native protocol; the gateway translates. This keeps backend services simple while supporting diverse client needs.
Benchmark the gateway with representative routes, payloads, TLS, and plugins at the target latency. Divide peak QPS by measured per-instance planning capacity, then add enough capacity for the largest failure you need to tolerate; two instances alone do not guarantee that either can carry peak traffic after the other fails.
Network bandwidth matters: outbound traffic is typically 8x inbound (responses are larger than requests). Memory sizing: ~1 GB per instance covers buffers, application heap, and OS overhead. Plan for failover capacity — during instance failure, remaining instances must handle full traffic.
A generic API gateway serves all client types through a single instance with unified routing and aggregation logic. A BFF is a specialized gateway instance tailored to a specific client type — mobile, web, or partner API — each with its own data requirements, payload shapes, and aggregation patterns.
Use BFF when mobile apps need smaller payloads with different fields than web apps, when teams have clear ownership boundaries (mobile team vs. web team), and when you have enough traffic to justify separate deployments. It adds complexity but gives per-client teams full autonomy over their gateway logic.
Rate limiting enforces a hard cap on the number of requests a client can make in a time window — excess requests get rejected with 429. Throttling smooths out traffic by queuing or slowing requests rather than dropping them outright.
At the gateway layer, rate limiting is the primary mechanism — it is simple to implement and gives clear pass/fail signals. Throttling is less common at the gateway because queued requests still hold gateway resources. Some gateways implement "delayed rejection" throttling where requests wait briefly before being rejected.
TLS termination at the gateway (edge termination) is the standard approach: clients terminate TLS at the gateway, and the gateway communicates with backend services over internal plaintext or mTLS. This reduces cryptographic overhead at scale and centralizes certificate management.
Re-encrypting for backend calls (mTLS between gateway and services) adds security for sensitive traffic but increases CPU overhead. For low-security internal networks, plaintext backend communication is acceptable as long as network isolation prevents direct access to services.
Full end-to-end TLS (client to backend, gateway as pass-through) adds maximum security but eliminates the gateway's ability to inspect, transform, or log request/response content.
Request aggregation — the gateway calling multiple backend services and combining responses — is powerful for client convenience but can cause memory pressure when responses are large or many services are called in parallel.
Memory issues arise when: a single aggregated response exceeds the gateway's memory limits, slow backend services cause the gateway to hold many in-flight responses simultaneously, or aggregation timeouts allow partial responses to accumulate.
Mitigations: set per-request memory limits, use streaming aggregation where possible, apply aggressive timeouts to individual backend calls, and cap the number of parallel backend calls the gateway will make for a single client request.
A service registry (e.g., Consul, etcd, Kubernetes endpoints) maintains the current list of healthy instances for each backend service. The gateway queries the registry to route requests, rather than using static configuration.
Health checks keep the registry accurate: the gateway or a separate process periodically calls each service instance's health endpoint and deregisters instances that fail. This ensures routing stops going to instances that are starting up, overloaded, or crashed.
Stale registry data is a common failure mode — instances can be dead but still in the registry if health checks are infrequent or the deregistration signal is missed. Use short TTLs (30 seconds or less) and ensure deregistration is event-driven, not just TTL-based.
To isolate gateway latency: measure time-to-first-byte at the gateway (before forwarding to backend) vs. backend response time. The gateway's internal processing time (auth, rate limiting, routing) should be captured as a separate histogram bucket.
Key metrics: gateway_request_duration_seconds (with backend_service and route labels), gateway_backend_latency_seconds (time spent waiting for backend), and gateway_overhead_seconds (calculated as total minus backend time).
Percentiles matter more than averages: P50 can look fine while P99 reveals latency tails caused by connection pool exhaustion, GC pauses, or slow rate limiting stores. Always look at P95 and P99 when diagnosing latency issues.
A gateway can be both a DDoS target and a DDoS shield. As a target, attackers aim traffic at the gateway to exhaust its resources. As a shield, the gateway's rate limiting and connection management can absorb or deflect attack traffic before it reaches backend services.
Protective measures at the gateway: aggressive rate limiting by IP and API key, connection limits per client, request size limits to prevent amplification, and IP blocklists for known bad actors. For volumetric DDoS (Gbps+ attacks), these are insufficient — edge DDoS protection (Cloudflare, AWS Shield) is needed before traffic reaches the gateway.
The gateway should also emit rate limit violation metrics so security teams can detect and respond to attack patterns in real time.
Configuration drift — the gateway behaving differently across environments due to subtle config differences — is a common operational problem. Rate limiting thresholds, routing rules, and feature flags often vary between environments in ways that cause prod-only bugs.
Best practices: store gateway config in version control with environment-specific overrides; use canary deployments for config changes (roll out to 5% of traffic first); treat config as code with code review requirements; and have automated config validation that runs before applying changes.
Secrets management is separate from config: use a secrets manager (Vault, AWS Secrets Manager) for API keys and credentials, not the gateway config file itself.
The gateway must identify tenants from incoming requests — via API key header, JWT claim, or subdomain — and ensure requests route only to that tenant's backend services. Tenant isolation is enforced at the routing layer, not left to backend services alone.
For shared backend services serving multiple tenants, the gateway should inject tenant context into request headers (X-Tenant-ID) so backend services can scope data queries. The gateway itself must never cache responses across tenants — a cached response for one tenant must not be served to another.
Rate limiting must be per-tenant, not global. A single misbehaving tenant should not consume budget that affects other tenants on the same gateway.
The gateway is the central enforcement point for API lifecycle management. It should add Deprecation and Sunset headers to responses for deprecated endpoints, track usage of deprecated API versions, and eventually block requests to sunset endpoints with clear migration guidance.
Deprecation workflow: announce deprecation 6+ months before sunset, add Deprecation: true and Sunset:
At sunset date, the gateway should return 410 Gone for deleted endpoints rather than generic 404, with a response body explaining the replacement version and migration steps.
Without graceful shutdown, a gateway instance being terminated loses in-flight requests — clients see connection errors mid-request. For a gateway handling hundreds or thousands of concurrent requests, this causes a spike of failed requests at every deployment.
Graceful shutdown involves: stopping new connections (draining the load balancer target), waiting for in-flight requests to complete (with a timeout), then exiting. Typical configuration: SIGTERM triggers graceful shutdown, 30-second timeout for in-flight requests, then force-kill if needed.
Health checks at /health can report unhealthy during drain, causing the load balancer to stop routing new traffic while existing requests complete.
The gateway timeout should always be shorter than the backend service timeout. If the gateway waits longer than the backend, the gateway times out first and returns an error for a request that might have succeeded — the backend wasted resources processing it.
Best practice: gateway timeout = backend timeout minus headroom for gateway processing (e.g., backend has 10s timeout, gateway uses 8s). This ensures the gateway returns a clean 504 before the backend sends a response the gateway will drop.
Different routes can have different timeout values based on the backend service's characteristics. Slow endpoints (report generation) get longer timeouts; fast endpoints (health checks) get shorter ones.
The gateway can implement retries for idempotent GET requests or those with explicit idempotency keys. Retries should use exponential backoff with jitter to avoid thundering herd problems. The gateway should add X-Request-ID to track retry chains.
Retries must be avoided for non-idempotent mutations (POST, DELETE without idempotency keys) as they can cause duplicate operations. POST to create an order should not be retried automatically — if it times out, the client should check order status before retrying.
Retries amplify failures: a backend at 50% capacity receiving retries goes to 100% and fails more. Circuit breakers should trip before retries amplify a degraded backend into a cascading failure.
Further Reading
Internal Resources
- Service Mesh — Managing internal service-to-service communication
- Load Balancing — Distributing traffic across multiple gateway instances
- RESTful API Design — Best practices for API contract design
- Circuit Breaker Pattern — Preventing cascade failures
- System Design Roadmap — Complete learning path for system design
External Resources
- NGINX API Gateway documentation — Production-grade reverse proxy and gateway setup
- Kong Gateway docs — Open-source API gateway with plugin ecosystem
- AWS API Gateway developer guide — Managed gateway on AWS with Lambda integration
- OAuth 2.0 RFC 6749 — The specification behind modern API authentication
- Rate Limiting Algorithms — Cloudflare blog — Deep dive on sliding window and token bucket at scale
Conclusion
An API gateway is the foundational piece that ties together client requests, backend services, and operational concerns like authentication, rate limiting, and observability. It simplifies client code by providing a single entry point, centralizes cross-cutting concerns so individual services stay thin, and gives you a central vantage point for monitoring, security, and traffic management.
The key decisions when adopting an API gateway are: choosing between a managed service or self-hosted solution, implementing Redis-backed rate limiting for consistent enforcement across instances, adding circuit breakers to prevent cascade failures, and evaluating whether a BFF pattern is needed for multi-client architectures.
Most production deployments require at least two gateway instances behind a load balancer, TLS termination at the edge, short JWT TTLs with refresh token rotation, and automated certificate renewal. Treat the gateway as a stateless proxy — keep business logic in backend services and store session state externally in Redis.
For most microservices architectures, an API gateway handles external client traffic while a service mesh handles internal service-to-service communication. Together they provide comprehensive coverage for north-south and east-west traffic patterns.
Category
Related Posts
Microservices vs Monolith: Choosing the Right Architecture
Understand the fundamental differences between monolithic and microservices architectures, their trade-offs, and how to decide which approach fits your project.
Server-Side Discovery: Load Balancer-Based Service Routing
Learn how server-side discovery uses load balancers and reverse proxies to route service requests in microservices architectures.
Amazon Architecture: Lessons from the Pioneer of Microservices
Learn how Amazon pioneered service-oriented architecture, the famous 'two-pizza team' rule, and how they built the foundation for AWS.