05 · Load Balancing¶
A load balancer sits in front of a pool of servers and decides which one receives each request or connection. It is what turns "we have ten app servers" into one service with one address. It also quietly handles failure: when a server dies, the balancer stops sending it traffic.
Layer 4 vs layer 7¶
L4 (transport-level) balancers see TCP/UDP connections: source and destination IPs and ports. They forward packets or connections without reading the HTTP inside.
- Very fast and cheap per connection; protocol-agnostic (databases, custom protocols).
- Cannot route by URL path or header, cannot retry a failed HTTP request, and balance per connection, not per request. With long-lived HTTP/2 or gRPC connections, one connection can carry many requests, so L4 balancing can be very uneven.
L7 (application-level) balancers terminate the connection, parse HTTP, and open their own connections to backends.
- Can route
/api/*to one pool and/static/*to another, route by header or cookie, terminate TLS, add headers, retry idempotent requests, and balance per request. - Cost more CPU, and they see decrypted traffic, which has security implications.
Most web architectures use L7 balancing (reverse proxies such as NGINX, HAProxy, Envoy, or cloud-managed equivalents) at the front, sometimes with L4 balancing in front of them for raw scale.
Balancing algorithms¶
| Algorithm | How it picks | Good for | Weak when |
|---|---|---|---|
| Round robin | Next server in order | Equal servers, similar requests | Requests vary widely in cost |
| Weighted round robin | Proportional to weight | Mixed machine sizes | Weights go stale |
| Least connections | Fewest open connections | Variable request duration | Many balancers each see partial counts |
| Least response time / EWMA | Fastest recently | Heterogeneous latency | Needs good measurements |
| Power of two choices | Pick 2 at random, choose the less loaded | Large pools, many balancers | — generally robust |
| Hash (IP, user ID, key) | Same key → same server | Session affinity, cache locality | Uneven keys; rehash on pool change |
Worked example: comparing algorithms in a simulation¶
This simulation sends requests with widely varying costs to five servers and measures the worst queue that builds up. It runs with the standard library.
# lb_sim.py — compare balancing strategies under uneven request costs
import random
def simulate(pick, servers=5, requests=20_000, seed=7):
rng = random.Random(seed)
busy_until = [0.0] * servers # time each server becomes free
worst_wait = 0.0
t = 0.0
rr = 0
for _ in range(requests):
t += rng.expovariate(4.5) # arrivals: ~4.5 per time unit
cost = rng.choice([0.1] * 9 + [8.0]) # 90% cheap, 10% very expensive
backlog = [max(0.0, b - t) for b in busy_until]
i, rr = pick(backlog, rr, rng)
start = max(t, busy_until[i])
worst_wait = max(worst_wait, start - t)
busy_until[i] = start + cost
return round(worst_wait, 1)
def round_robin(backlog, rr, rng):
return rr % len(backlog), rr + 1
def least_loaded(backlog, rr, rng):
return min(range(len(backlog)), key=backlog.__getitem__), rr
def two_choices(backlog, rr, rng):
a, b = rng.sample(range(len(backlog)), 2)
return (a if backlog[a] <= backlog[b] else b), rr
for name, fn in [("round robin", round_robin),
("least loaded", least_loaded),
("power of two", two_choices)]:
print(f"{name:13s} worst wait: {simulate(fn)}")
Run it and you should see round robin produce a much larger worst-case wait than the two load-aware strategies — typically several times larger — while least-loaded and power-of-two land in the same range (sometimes one wins, sometimes the other). Exact numbers depend on the random seed, which is why the lesson does not quote them — change the seed and the cost mix and observe that the ordering holds. Round robin blindly queues cheap requests behind expensive ones; the load-aware strategies route around busy servers. (Here "least loaded" uses perfect knowledge of every backlog, which a real balancer does not have; it approximates it with connection counts or latency.)
Health checks and draining¶
A balancer must know which backends are able to serve.
- Active health checks: the balancer calls
GET /healthzevery few seconds. After a few consecutive failures the backend is marked down; after a few successes, up again. - Passive health checks (outlier detection): the balancer notices real requests failing or timing out and ejects the backend temporarily.
A good health endpoint checks that the process can do useful work (for instance, that it can reach its required dependencies without making the health check itself expensive). Be careful: if every server's health check fails because a shared database is down, the balancer removes all servers and turns a partial outage into a total one. Many teams separate a liveness check (is the process alive?) from a readiness check (should it receive traffic?).
Connection draining: during deploys, a server is marked as leaving; the balancer stops sending new requests but lets in-flight ones finish before the server shuts down. Without this, every deploy produces a burst of errors.
Removing the balancer as a single point of failure¶
A single load balancer box is itself a single point of failure. Common approaches:
- Active-passive pair sharing a virtual IP; the standby takes over the IP if the active one stops sending heartbeats.
- Several active balancers behind DNS (multiple A records) or anycast routing, so the same IP is announced from several places.
- Managed cloud balancers, which are internally distributed; you still depend on that provider's regional availability.
flowchart TB
DNS[DNS: api.example.com] --> LB1[LB 1]
DNS --> LB2[LB 2]
LB1 --> S1[App 1]
LB1 --> S2[App 2]
LB1 --> S3[App 3]
LB2 --> S1
LB2 --> S2
LB2 --> S3
How It Actually Works¶
An L7 balancer is a reverse proxy running an event loop. It accepts the client's TCP
connection, completes TLS using the site's certificate, and parses the HTTP request.
It then picks a backend using its algorithm and its current view of backend health, and
either reuses an idle pooled connection to that backend or opens one. It forwards the
request (often adding headers such as X-Forwarded-For so the backend knows the
client's IP), streams the response back, and returns the backend connection to the pool.
Because several balancer instances run at once, each sees only its own traffic. Their "least connections" counts are local, not global. That is one reason the power of two choices is popular: choosing the better of two random backends gives most of the benefit of global least-loaded with no coordination at all, and it avoids the "herd" effect where every balancer sends traffic to the same apparently idle server at once.
L4 balancers work lower down. Some rewrite packet addresses (NAT) and see return traffic; others use direct server return, where the backend replies to the client directly, bypassing the balancer for the (usually much larger) response. They typically choose a backend by hashing the connection's addresses and ports, so all packets of one connection reach the same server.
Common mistakes¶
- Sticky sessions as a crutch. Pinning users to servers hides stateful app servers; it breaks when that server dies and skews load.
- Retrying non-idempotent requests at the balancer, causing duplicate writes.
- Health checks that are too deep (take down everything when a dependency blips) or too shallow (report healthy while the app is deadlocked).
- Forgetting the balancer's own capacity and connection limits during peaks.
- L4 balancing for gRPC/HTTP/2 and wondering why one backend is overloaded.
Exercise¶
- Run
lb_sim.pywith five different seeds and two different cost mixes. Record which algorithm wins each time and explain why. - Add a "hash by client ID" strategy with 1,000 simulated clients and compare. What happens if one client sends 20% of all traffic?
- Design the health-check policy for an app that depends on a database and an optional recommendations service. Which failures should remove a server from rotation, and which should not?