Level 3 · Advanced Distributed correctness¶
Levels 1 and 2 were about capacity: how to serve more requests and store more data. Level 3 is about correctness when things go wrong — which, at scale, is all the time. Nodes crash mid-operation, messages arrive twice, clocks disagree, and a slow dependency drags down everything that calls it.
These lessons cover the mechanisms that let systems agree, coordinate, and recover: consensus for a single source of truth, sagas for workflows spanning services, idempotency for surviving retries, event streams for keeping systems in sync, and the observability and reliability patterns that let you run all of it.
Modules¶
- Consensus Basics: Raft Conceptually — leader election, log replication, terms, and why majorities matter
- Distributed Transactions & Sagas — two-phase commit, its blocking problem, and saga compensation
- Idempotency & the Exactly-Once Myth — idempotency keys, dedup windows, and what "exactly once" really means
- Event-Driven Architecture & CDC — events vs commands, the outbox pattern, and change data capture
- Search Systems — inverted indexes, relevance scoring, indexing pipelines, and sharded search
- Real-Time Systems: WebSockets & Pub/Sub — polling vs SSE vs WebSockets, connection gateways, and presence
- Observability — metrics, logs, traces, SLOs, and error budgets
- Reliability Patterns — timeouts, retries with backoff and jitter, circuit breakers, bulkheads, load shedding
- Hot Keys, Skew & Multi-Tenancy — detecting and handling uneven load and noisy neighbours
- Project — Design a Chat System — one-to-one and group messaging, delivery guarantees, ordering, and presence
What you need before starting¶
- Levels 1–2, particularly replication, partitioning, queues, and CAP/PACELC.
- Python 3 for the simulations (a toy Raft election, a saga orchestrator, a circuit breaker). As before, only the standard library is required.
This is the densest level. Expect to reread lessons 1–3; they are the foundation of every serious distributed-systems discussion.