At a glance
Key takeaways
- Autonomy spectrum: fixed workflow → router → agent loop → multi-agent. Use the least that works.
- Named patterns: chaining (with gates), routing (with a fallback), parallelization (sectioning, voting), orchestrator-workers, evaluator-optimizer.
- Multi-agent buys focused context and parallelism, and costs tokens, hand-off losses and debuggability.
- Production shape: an explicit state machine, the model consulted only at specific nodes, checkpoints after each step.
Level 2
How it works, from scratch
"Agent" covers everything from one model call inside ordinary code to a team of models steering each other. This lesson lays out that range, builds each named pattern in a few dozen lines, and measures what each step up costs. The rule that falls out is simple: use the least autonomy that solves the problem, and say so out loud.
Chapter 1
The autonomy spectrum
Everyday picture Four ways to run an office. An assembly line (fixed workflow): the steps never change and people do one task at each station. A receptionist (router) decides which department you go to. A personal assistant (agent loop) decides which errands to run and when the job's done. A team with a manager (multi-agent) splits the work among specialists.
Figure 1 · Diagram
flowchart LR W[Fixed workflow<br/>code decides steps] --> R[Router<br/>LLM picks a path] R --> A[Agent loop<br/>LLM picks tools] A --> M[Multi-agent<br/>agents coordinate]
Tiny worked example One question, "How many PTO days do I get, and
what is the meal per-diem when travelling?", answered by each design
(autonomy_costs()):
| design | model calls | why |
|---|---|---|
| fixed workflow | 2 | rewrite as a search, then answer |
| router | 2 | classify, then the chosen handler answers |
| agent loop | 2 | one turn with two parallel tool calls, one answer turn |
| multi-agent | 4 | plan, one call per specialist (2), final answer |
Figure 2 · Drawn from the lesson's code
Workflow, router and agent loop each make 2 calls and multi-agent 4, but the agent loop uses about 990 tokens, ten times the workflow's 93
Chapter 2
The named patterns
These names (from Anthropic's Building effective agents) let you describe a design in one word.
In code: every pattern below is built from ask, one model call with
one user message that returns the reply's text.
Prompt chaining
Everyday picture A relay race with an inspector at each hand-off: if the baton is dropped, the next runner doesn't start.
Tiny worked example Step 1 writes an outline, and a code gate checks it has at
least three points. The outline "- scope" fails the gate, so the chain stops after one
call and the draft step never runs.
Figure 3 · Diagram
flowchart LR
I[Input] --> S1[LLM: outline] --> G{Gate: 3+ points?}
G -->|yes| S2[LLM: draft] --> O[Output]
G -->|no| X[Stop with the reason]
In code: run_chain runs a list of ChainSteps in order, feeding each
output into the next prompt and stopping at the first gate that reports a
problem; ChainResult records the outputs and where and why it stopped.
Routing
Everyday picture The receptionist listens for five seconds and sends you to billing, tech support or the front desk.
Tiny worked example The classifier replies " Technical\n", which is
normalized to technical and routed. It replies "Billing-ish, maybe refunds?",
which is not a known label, so the request goes to the general fallback
instead of crashing.
Figure 4 · Diagram
flowchart LR Q[Request] --> C[LLM classifier] C -->|billing| B[Billing handler] C -->|technical| T[Tech handler] C -->|anything else| F[Fallback]
In code: route makes the one classifier call, normalizes the label,
swaps anything unknown for the fallback, and hands the request to that
label's handler.
Parallelization: sectioning and voting
Everyday picture Sectioning: several cooks each make one dish at the same time. Voting: a panel of judges, where the majority decides.
Tiny worked example Three reviewers judge a SQL snippet: two say
"vulnerable" and one says "safe", so the verdict is ("vulnerable", {"vulnerable": 2, "safe": 1}).
Figure 5 · Diagram
flowchart LR
Q[Same prompt] --> R1[Reviewer 1]
Q --> R2[Reviewer 2]
Q --> R3[Reviewer 3]
R1 --> V{Majority}
R2 --> V
R3 --> V
Why voting helps, when voters err independently:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning |
|---|---|
| number of voters (odd, so there are no ties) | |
| chance a single voter is right | |
| number of voters who are right | |
| ways to choose which of the are right |
In words: add up the chances of every outcome where more than half the voters are right.
On the example: , : . Three 80% judges make an 89.6% panel.
In Python:
from math import comb
n, p = 3, 0.8
# the chance exactly k voters are right ...
P = sum(comb(n, k) * p ** k * (1 - p) ** (n - k)
# ... for every k from ⌊n/2⌋ + 1 to n
for k in range(n // 2 + 1, n + 1))
round(P, 3) # → 0.896
Figure 6 · Drawn from the lesson's code
Majority accuracy rises with more independent voters: at 70% per voter, 5 voters reach 84% and 15 reach 95%; at 60% per voter the climb is slow
In code: run_sections is sectioning: it runs independent pieces of
work on a thread pool and collects every result. vote is voting: it asks
each reviewer the same prompt through run_sections and returns the
majority label with its tally, and majority_accuracy evaluates the formula
above.
Orchestrator-workers
Everyday picture A project lead reads the brief, splits it into sections, hands each to a writer, then edits the pieces into one report.
Tiny worked example "What's the meal per-diem, when are receipts due,
and does PTO roll over?" The orchestrator returns three subtasks. Three workers each
find one document (fin-002, fin-004, hr-001), and the synthesizer
writes one answer citing all three.
Figure 7 · Diagram
flowchart TD Q[Request] --> O[Orchestrator LLM:<br/>decide the subtasks] O --> W1[Worker: per-diem] O --> W2[Worker: receipts] O --> W3[Worker: PTO rollover] W1 --> S[Synthesizer LLM:<br/>one cited answer] W2 --> S W3 --> S
In code: orchestrate asks the orchestrator for subtasks, runs the
workers on them in parallel with run_sections, and asks the synthesizer for
one answer, returning an OrchestratorResult. policy_worker is the worker
in the example: it finds the one current policy document for a subtask.
Evaluator-optimizer
Everyday picture A writer and an editor. The draft goes back and forth until the editor signs off, or the deadline hits.
Tiny worked example Draft 1, "PTO rolls over.", gets the feedback "Missing a citation". Draft 2, "Up to 5 unused PTO days roll over [hr-001].", passes. Result: 2 rounds, passed.
Figure 8 · Diagram
flowchart LR
T[Task] --> G[Generator drafts]
G --> E{Evaluator:<br/>PASS?}
E -->|feedback| G
E -->|PASS| D[Done]
E -->|max rounds| B[Stop: best effort, flagged]
In code: evaluate_optimize alternates generator drafts and evaluator
verdicts, feeding each critique back, and returns the last draft, the rounds
used and whether it passed.
Chapter 3
Multi-agent: a supervisor and specialists
Everyday picture A manager who never does the work, only assigns it and assembles the results. Each specialist gets a short, focused brief.
Tiny worked example The supervisor sends "PTO days per year?" to hr and
"Meal per-diem?" to finance. The HR specialist never sees the finance question,
and vice versa (tested). A delegation to a non-existent legal specialist
is reported, not crashed on.
Figure 9 · Diagram
sequenceDiagram participant U as User participant S as Supervisor participant H as HR agent participant F as Finance agent U->>S: PTO and meal allowance? S->>H: "PTO days per year?" (only this) S->>F: "Meal per-diem?" (only this) H-->>S: 20 days [hr-001] F-->>S: 75 dollars a day [fin-002] S-->>U: combined answer
When it's worth it. Use several agents when subtasks are genuinely separable, such as independent research threads or per-document work that would overflow one context. The costs are real: more tokens, context lost at every hand-off, and much harder debugging.
In code: supervise gets a delegation plan from the supervisor, sends
each specialist only its own task (reporting unknown names instead of
crashing), then asks the supervisor for the final answer; SupervisorResult
keeps the delegations and every specialist's answer.
Chapter 4
An explicit state machine, with the model at specific nodes
Everyday picture A board game. The squares and the rules for moving between them are fixed. On a few special squares you draw a card (ask the model). Everywhere else, the rules decide.
Tiny worked example Invoice intake. The model is asked twice: "is this an invoice?" and "extract the number, vendor and amount". Code does everything else: validation, the approval rule (over 5,000 waits for a person), and posting. The runs:
INVOICE INV-2041 ... 1200.00 -> RECEIVED > CLASSIFIED > EXTRACTED > VALIDATED > POSTED > DONE
INVOICE INV-2042 ... 18000.00 -> RECEIVED > CLASSIFIED > EXTRACTED > VALIDATED > AWAITING_APPROVAL
Acme newsletter -> RECEIVED > REJECTED
Figure 10 · Diagram
stateDiagram-v2 [*] --> RECEIVED RECEIVED --> CLASSIFIED: LLM says invoice RECEIVED --> REJECTED: LLM says other CLASSIFIED --> EXTRACTED: LLM extracts fields EXTRACTED --> VALIDATED: code checks pass EXTRACTED --> NEEDS_REVIEW: code checks fail VALIDATED --> AWAITING_APPROVAL: amount > limit VALIDATED --> POSTED: code posts to ledger AWAITING_APPROVAL --> POSTED: human approves POSTED --> DONE
In code: InvoiceWorkflow.run walks the states, calling the model only
at RECEIVED and CLASSIFIED; validate_invoice is the code check between
EXTRACTED and VALIDATED; InvoiceWorkflow.approve is the human step that
releases a paused invoice. WorkflowRun holds the current state, its history
and the extracted data.
Checkpoints and durable execution. After every transition the run is saved to disk. Durable execution means a workflow whose progress survives crashes, because each completed step's result is stored and the run resumes from there. Engines like Temporal do this at scale.
Figure 11 · Drawn from the lesson's code
After one crash while posting, resuming from the checkpoint finishes in 2 model calls, while restarting from scratch takes 4
In code: InvoiceWorkflow saves a checkpoint file after every
transition and loads it at the start of InvoiceWorkflow.run, so a rerun
resumes. resume_comparison crashes the ledger once and counts the model
calls each way.
Chapter 5
Frameworks vs. plain code
LangGraph models agents as graphs with explicit state and checkpoints. The Claude Agent SDK and OpenAI Agents SDK give you vendor-built agent loops, tools and hand-offs. CrewAI and AutoGen focus on multi-agent teams. Temporal provides durable execution for any code. A framework earns its place with standard patterns, built-in tracing, persistence and human-in-the-loop hooks. Plain code wins for simple flows, full control and fewer dependencies. Everything in this lesson is under 100 lines of plain Python, and knowing that is what lets you judge a framework.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Q: When would you use a fixed workflow instead of an agent? Give an example.Think it through, then reveal
A: When the steps are known in advance and the variation is inside each step, not in which steps happen. Invoice intake is the example: classify, extract, validate, approve, post. A workflow is cheaper, faster, testable and auditable. The model is used where judgement is needed (classification, extraction), and code owns the flow. Reach for an agent only when the path genuinely depends on what's discovered along the way.
Question 2Q: Single agent vs. multi-agent for a document-processing workflow: what does each side have going for it?Think it through, then reveal
A: For one agent (or a workflow): documents flow through the same steps, hand-offs lose context, multi-agent multiplies tokens, and one trace is far easier to debug. For several agents: documents are independent and can be processed in parallel, each document or section may overflow a single context, and specialists (tables, legal clauses, figures) benefit from focused instructions and tools. A common answer is a workflow that fans out per document to a focused worker, which is orchestrator-workers rather than free-form agents talking to each other.
Question 3Q: A router's classifier sometimes returns labels that aren't in its list. How should the code handle that?Think it through, then reveal
A: Normalize (case, whitespace), accept only known labels, and send
everything else to a safe fallback. Log the misses and add them to the
classifier's evaluation set. Consider structured output with an enum so
the model can only produce valid labels.
Question 4Q: Why checkpoint after every step of a long workflow?Think it through, then reveal
A: So a crash costs one step, not the run. Resuming reuses completed work (no repeated model calls, no double-posted invoices when combined with idempotency keys) and avoids getting a different answer on the rerun.
Primary sources
The papers behind this lesson
It established the reason-act-observe loop that the "agent loop" rung of the spectrum is built on.
Read the annotated companion →The paper ↗Researcher's shelf
Further reading
- Anthropic, Building effective agents (the named patterns): https://www.anthropic.com/engineering/building-effective-agents
- LangGraph docs: https://langchain-ai.github.io/langgraph/
- OpenAI Agents SDK: https://openai.github.io/openai-agents-python/
- CrewAI docs: https://docs.crewai.com/
- AutoGen: https://microsoft.github.io/autogen/
- Temporal, durable execution: https://docs.temporal.io/
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.