Level 3 · Advanced Scale & Rigor¶
By now you can build an agent that works on the examples you tried. Level 3 is about the gap between "works on my examples" and "works, measurably, on the inputs we will actually get, at a cost we can afford, without being turned against us".
The level has three threads:
- Composition — splitting work across several agents (lessons 01–02), and connecting agents to tools through a standard protocol, MCP (lesson 03). Plus the special case of agents that write and run code, which demands real isolation (lesson 04).
- Measurement — evaluating agents by outcome and by trajectory (lessons 05–06), then using those numbers to control cost and latency (lesson 07).
- Robustness — retries, timeouts and idempotency so that failures don't double-charge anyone (lesson 08), and defending against prompt injection arriving through tools and data (lesson 09).
The project (lesson 10) builds a small multi-agent triage system and, more importantly, the evaluation harness that tells you whether it's any good.
Modules¶
- Multi-Agent Systems: Why and When — what splitting buys you, what it costs, and a test for whether you need it
- Supervisors & Handoffs — two coordination patterns, implemented and compared
- Model Context Protocol (MCP) Concepts — hosts, clients, servers, primitives and the trust model
- Sandboxing Code-Execution Agents — timeouts, limits, isolation layers, and what a subprocess can't protect
- Evaluating Agents: Task Success — datasets, checkers, repeated trials and pass rates
- Trajectory Evaluation — judging how the agent got there: required, forbidden and wasted steps
- Cost & Latency Control — token accounting, routing, caching and parallel tool calls
- Reliability: Retries, Timeouts, Idempotency — making actions safe to repeat
- Prompt Injection via Tools & Data — how untrusted content steers agents, and layered defenses
- Project — Evaluated Multi-Agent Triage — a supervisor, two specialists and an eval report
Before you start¶
- The code continues from Levels 1 and 2 (
tools.py,mini_agent.py,guards.py,tracer.py,approval.py). Python 3.9+, standard library only. - Lesson 04 runs code in a subprocess and uses the Unix-only
resourcemodule; on Windows those limits are skipped and the lesson explains the alternatives. - Security material in this level is defensive: how attacks on agents work in principle and how to design so they fail. For broader application security, see the Cybersecurity Mastery Path; for LLM-level safety and red-teaming, see the LLM Dev Mastery Path.