03 · API Gateways Basics¶
An API gateway sits in front of your services as the single entry point clients talk to. Instead of every client knowing the address of every backend service, they all hit the gateway, which routes, secures, and shapes traffic on their behalf.
What a gateway actually does¶
- Routing —
/v1/orders/*→ orders-service,/v1/users/*→ users-service, based on path, host, or header. - Authentication — validates the token once, at the edge, so individual services don't each re-implement JWT verification.
- Rate limiting — enforces per-client quotas centrally (see module 9 of Level 2).
- TLS termination — the gateway holds the certificate; internal traffic can run plain HTTP inside a trusted network.
- Request/response transformation — rewriting headers, aggregating multiple backend calls into one client-facing response.
- Observability — one place to log every request, emit metrics, and trace latency across the whole surface.
A minimal gateway config (Kong-style, conceptually)¶
services:
- name: orders-service
url: http://orders.internal:8080
routes:
- name: orders-route
paths: ["/v1/orders"]
plugins:
- name: rate-limiting
config: { minute: 100 }
- name: jwt
A request from a client never sees orders.internal — it only ever
talks to https://api.example.com/v1/orders, and the gateway resolves
where that actually goes.
Routing in action¶
The gateway:
- Terminates TLS.
- Validates the JWT signature and expiry.
- Checks the rate-limit bucket for this client.
- Matches
/v1/orders/*to theorders-serviceroute. - Proxies the request internally, adding a trace header:
GET /orders/42 HTTP/1.1
Host: orders.internal:8080
X-Request-Id: 7f3a-91c2-...
X-Forwarded-For: 203.0.113.9
- Streams the backend's response back to the client unchanged (or transformed, per config).
Gateway vs. load balancer vs. service mesh¶
| Layer | Operates on | Typical job |
|---|---|---|
| Load balancer | L4/L7, IPs and ports | Spread traffic across replicas |
| API gateway | L7, HTTP semantics | Auth, routing, rate limiting, per-API concerns, client-facing |
| Service mesh (e.g. Istio, Linkerd) | L7, service-to-service | mTLS, retries, circuit breaking between internal services |
They're complementary, not competing: a gateway handles north-south traffic (client → cluster); a mesh handles east-west traffic (service → service) once the request is inside.
Aggregation: one client call, many backend calls¶
A mobile app's dashboard screen needs data from three services. Instead of the client making three round-trips, the gateway (or a backend-for-frontend layer behind it) fans out internally:
{
"profile": { "id": 42, "name": "Ada" },
"orders": { "open_count": 2 },
"notifications": { "unread": 5 }
}
Internally this is three parallel calls to users-service,
orders-service, and notifications-service, merged before responding
— saving mobile clients three round-trips over a slow network.
Worked example: adding a new service behind the gateway¶
Your team ships a new reviews-service. Rollout without breaking
anything:
- Deploy
reviews-serviceinternally, unreachable from outside the cluster. - Add a gateway route:
/v1/reviews/*→reviews-service, with the same JWT and rate-limit plugins every other route uses. - Deploy behind a feature flag or to a small percentage of traffic first (gateway-level canary routing) if the platform supports it.
- Only after the gateway route is live does
reviews-servicebecome reachable by clients at all — until then it's fully internal, so there's zero external-facing risk during development.
How It Actually Works¶
An API gateway sits as a reverse proxy in front of your actual services, and its core mechanism is inspecting and rewriting each request before forwarding it — a single incoming connection from the client, then a separate outgoing connection to the real backend.
Client -> [Gateway] -> Backend service
|
+- terminates TLS from the client
+- matches path against a routing table
+- runs auth/rate-limit checks BEFORE forwarding
+- opens a NEW connection to the matched backend
+- (often) rewrites the path, e.g. strips /api prefix
+- streams the backend's response back to the client
The gateway terminating TLS means the client's encrypted connection ends at the gateway — traffic from gateway to backend may be plain HTTP inside a trusted private network, which is why backend services often don't need their own TLS certificates at all.
Centralizing auth and rate limiting at the gateway means those checks run
exactly once, before the request reaches any backend service — a
short-circuited 401/429 response from the gateway never even opens a
connection to the backend, saving that service's resources entirely. This
is mechanically why gateways reduce backend code duplication: instead of
every microservice independently verifying a JWT, one gateway process does
it once per request and forwards a already-verified user identity (often
as an internal header like X-User-Id) that backends trust implicitly
because they only accept traffic from the gateway's internal network.
Exercise¶
- A client complains that two different backend services return inconsistent error formats for the same kind of failure (missing auth). How would centralizing auth at the gateway fix this?
- Why does TLS termination usually happen at the gateway rather than at every individual service?
- Explain the north-south vs. east-west distinction and give an example of a concern that belongs at each layer.
- Your gateway aggregates three backend calls into one response. What should happen if one of the three backend calls fails — return partial data, or fail the whole request? Justify your answer for a dashboard endpoint specifically.