rumblr Work in progressWIP

● The AI Primer · Lesson 39 · Part 2: building systems people rely on

Orchestration

how much autonomy to give the model

This lesson covers Workflows vs. agents, and the named patterns

Members · open during launch 21 min11 figures and diagrams
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Autonomy spectrum: fixed workflow → router → agent loop → multi-agent. Use the least that works.
  2. Named patterns: chaining (with gates), routing (with a fallback), parallelization (sectioning, voting), orchestrator-workers, evaluator-optimizer.
  3. Multi-agent buys focused context and parallelism, and costs tokens, hand-off losses and debuggability.
  4. Production shape: an explicit state machine, the model consulted only at specific nodes, checkpoints after each step.

Level 2

How it works, from scratch

"Agent" covers everything from one model call inside ordinary code to a team of models steering each other. This lesson lays out that range, builds each named pattern in a few dozen lines, and measures what each step up costs. The rule that falls out is simple: use the least autonomy that solves the problem, and say so out loud.

Chapter 1

The autonomy spectrum

Everyday picture Four ways to run an office. An assembly line (fixed workflow): the steps never change and people do one task at each station. A receptionist (router) decides which department you go to. A personal assistant (agent loop) decides which errands to run and when the job's done. A team with a manager (multi-agent) splits the work among specialists.

Figure 1 · Diagram

Reading it: left to right, the model decides more and your code decides less. Each arrow buys flexibility (the system can handle requests you didn't anticipate) and costs predictability, tokens and debuggability. The skill is stopping at the leftmost box that handles your real requests.

Tiny worked example One question, "How many PTO days do I get, and what is the meal per-diem when travelling?", answered by each design (autonomy_costs()):

design model calls why
fixed workflow 2 rewrite as a search, then answer
router 2 classify, then the chosen handler answers
agent loop 2 one turn with two parallel tool calls, one answer turn
multi-agent 4 plan, one call per specialist (2), final answer

Figure 2 · Chart

fixed workflow router agent loop multi-agent 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 model calls Calls for the same question fixed workflow router agent loop multi-agent 0 200 400 600 800 1000 tokens (input + output) Tokens for the same question

Workflow, router and agent loop each make 2 calls and multi-agent 4, but the agent loop uses about 990 tokens, ten times the workflow's 93

Reading it: the left panel counts model calls and the right counts tokens. Calls barely move until multi-agent doubles them. Tokens tell a sharper story. The agent loop re-sends its tool definitions and tool results on every call, so it uses many times the tokens of the fixed workflow for the same answer. The supervisor pays for planning and hand-offs. None of that is waste if you need the flexibility, and all of it is waste if you don't.

Chapter 2

The named patterns

These names (from Anthropic's Building effective agents) let you describe a design in one word.

In code: every pattern below is built from ask, one model call with one user message that returns the reply's text.

Prompt chaining

Everyday picture A relay race with an inspector at each hand-off: if the baton is dropped, the next runner doesn't start.

Tiny worked example Step 1 writes an outline, and a code gate checks it has at least three points. The outline "- scope" fails the gate, so the chain stops after one call and the draft step never runs.

Figure 3 · Diagram

Reading it: two model calls in a fixed order, with plain code in the middle. The gate is where you catch a bad intermediate result before paying for, and being misled by, the next step.

In code: run_chain runs a list of ChainSteps in order, feeding each output into the next prompt and stopping at the first gate that reports a problem; ChainResult records the outputs and where and why it stopped.

Routing

Everyday picture The receptionist listens for five seconds and sends you to billing, tech support or the front desk.

Tiny worked example The classifier replies " Technical\n", which is normalized to technical and routed. It replies "Billing-ish, maybe refunds?", which is not a known label, so the request goes to the general fallback instead of crashing.

Figure 4 · Diagram

Reading it: one cheap model call decides the path, then specialised handling takes over. The "anything else" arrow is essential, because the label is free text from a model and code must never assume it's one of the keys.

In code: route makes the one classifier call, normalizes the label, swaps anything unknown for the fallback, and hands the request to that label's handler.

Parallelization: sectioning and voting

Everyday picture Sectioning: several cooks each make one dish at the same time. Voting: a panel of judges, where the majority decides.

Tiny worked example Three reviewers judge a SQL snippet: two say "vulnerable" and one says "safe", so the verdict is ("vulnerable", {"vulnerable": 2, "safe": 1}).

Figure 5 · Diagram

Reading it: the three calls run at the same time (wall-clock time is the slowest one), and plain code counts the answers. Sectioning looks the same except each branch gets a different sub-task, such as a guardrail check running beside the answer.

Why voting helps, when voters err independently:

Level 3: the formula and its symbols

Symbols

Symbol Meaning
number of voters (odd, so there are no ties)
chance a single voter is right
number of voters who are right
ways to choose which of the are right

In words: add up the chances of every outcome where more than half the voters are right.

On the example: , : . Three 80% judges make an 89.6% panel.

In Python:

from math import comb
n, p = 3, 0.8
# the chance exactly k voters are right ...
P = sum(comb(n, k) * p ** k * (1 - p) ** (n - k)
        # ... for every k from ⌊n/2⌋ + 1 to n
        for k in range(n // 2 + 1, n + 1))
round(P, 3)  # → 0.896

Figure 6 · Chart

2 4 6 8 10 12 14 number of independent voters (odd) 0.5 0.6 0.7 0.8 0.9 1.0 P(majority is right) Voting: many noisy judges beat one (if their errors are independent) each voter right 60% each voter right 70% each voter right 80% each voter right 90%

Majority accuracy rises with more independent voters: at 70% per voter, 5 voters reach 84% and 15 reach 95%; at 60% per voter the climb is slow

Reading it: each curve is a per-voter accuracy, and the x-axis adds voters. Every curve rises: more independent judges, better majority. The catch is the word independent. Five copies of the same model with the same prompt tend to make the same mistake, and then voting buys little. Vary the prompt, the model or the evidence.

In code: run_sections is sectioning: it runs independent pieces of work on a thread pool and collects every result. vote is voting: it asks each reviewer the same prompt through run_sections and returns the majority label with its tally, and majority_accuracy evaluates the formula above.

Orchestrator-workers

Everyday picture A project lead reads the brief, splits it into sections, hands each to a writer, then edits the pieces into one report.

Tiny worked example "What's the meal per-diem, when are receipts due, and does PTO roll over?" The orchestrator returns three subtasks. Three workers each find one document (fin-002, fin-004, hr-001), and the synthesizer writes one answer citing all three.

Figure 7 · Diagram

Reading it: it looks like sectioning, but the subtasks aren't known in advance. The orchestrator decides them from the request, which is what makes it suited to open-ended requests.

In code: orchestrate asks the orchestrator for subtasks, runs the workers on them in parallel with run_sections, and asks the synthesizer for one answer, returning an OrchestratorResult. policy_worker is the worker in the example: it finds the one current policy document for a subtask.

Evaluator-optimizer

Everyday picture A writer and an editor. The draft goes back and forth until the editor signs off, or the deadline hits.

Tiny worked example Draft 1, "PTO rolls over.", gets the feedback "Missing a citation". Draft 2, "Up to 5 unused PTO days roll over [hr-001].", passes. Result: 2 rounds, passed.

Figure 8 · Diagram

Reading it: the loop only exits on PASS or on the round limit. Without the limit, a critic that's never satisfied loops forever. Clear, checkable criteria make this pattern work.

In code: evaluate_optimize alternates generator drafts and evaluator verdicts, feeding each critique back, and returns the last draft, the rounds used and whether it passed.

Chapter 3

Multi-agent: a supervisor and specialists

Everyday picture A manager who never does the work, only assigns it and assembles the results. Each specialist gets a short, focused brief.

Tiny worked example The supervisor sends "PTO days per year?" to hr and "Meal per-diem?" to finance. The HR specialist never sees the finance question, and vice versa (tested). A delegation to a non-existent legal specialist is reported, not crashed on.

Figure 9 · Diagram

Reading it: the supervisor's messages to each specialist are the hand-off, and they're all each specialist knows. That's both the benefit (small, focused context) and the risk (whatever the supervisor forgets to write down is lost). Peer designs, where agents hand control directly to each other, have the same hand-off problem without a central coordinator.

When it's worth it. Use several agents when subtasks are genuinely separable, such as independent research threads or per-document work that would overflow one context. The costs are real: more tokens, context lost at every hand-off, and much harder debugging.

In code: supervise gets a delegation plan from the supervisor, sends each specialist only its own task (reporting unknown names instead of crashing), then asks the supervisor for the final answer; SupervisorResult keeps the delegations and every specialist's answer.

Chapter 4

An explicit state machine, with the model at specific nodes

Everyday picture A board game. The squares and the rules for moving between them are fixed. On a few special squares you draw a card (ask the model). Everywhere else, the rules decide.

Tiny worked example Invoice intake. The model is asked twice: "is this an invoice?" and "extract the number, vendor and amount". Code does everything else: validation, the approval rule (over 5,000 waits for a person), and posting. The runs:

INVOICE INV-2041 ... 1200.00   -> RECEIVED > CLASSIFIED > EXTRACTED > VALIDATED > POSTED > DONE
INVOICE INV-2042 ... 18000.00  -> RECEIVED > CLASSIFIED > EXTRACTED > VALIDATED > AWAITING_APPROVAL
Acme newsletter                -> RECEIVED > REJECTED

Figure 10 · Diagram

Reading it: every box is a state and every arrow a transition, and only two arrows are labelled "LLM". The rest are rules you can read, test and audit. When something goes wrong, you know exactly which state it was in and why it moved.

In code: InvoiceWorkflow.run walks the states, calling the model only at RECEIVED and CLASSIFIED; validate_invoice is the code check between EXTRACTED and VALIDATED; InvoiceWorkflow.approve is the human step that releases a paused invoice. WorkflowRun holds the current state, its history and the extracted data.

Checkpoints and durable execution. After every transition the run is saved to disk. Durable execution means a workflow whose progress survives crashes, because each completed step's result is stored and the run resumes from there. Engines like Temporal do this at scale.

Figure 11 · Chart

resume from checkpoint restart from scratch 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 model calls to finish one invoice Posting crashed once: what did it cost? 2 4

After one crash while posting, resuming from the checkpoint finishes in 2 model calls, while restarting from scratch takes 4

Reading it: the same crash happens in both bars. Resuming from the checkpoint finishes with the 2 model calls already made. Restarting from scratch asks the model both questions again, doubling cost, and, worse, risks a different answer the second time.

In code: InvoiceWorkflow saves a checkpoint file after every transition and loads it at the start of InvoiceWorkflow.run, so a rerun resumes. resume_comparison crashes the ledger once and counts the model calls each way.

Chapter 5

Frameworks vs. plain code

LangGraph models agents as graphs with explicit state and checkpoints. The Claude Agent SDK and OpenAI Agents SDK give you vendor-built agent loops, tools and hand-offs. CrewAI and AutoGen focus on multi-agent teams. Temporal provides durable execution for any code. A framework earns its place with standard patterns, built-in tracing, persistence and human-in-the-loop hooks. Plain code wins for simple flows, full control and fewer dependencies. Everything in this lesson is under 100 lines of plain Python, and knowing that is what lets you judge a framework.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: When would you use a fixed workflow instead of an agent? Give an example.Think it through, then reveal

A: When the steps are known in advance and the variation is inside each step, not in which steps happen. Invoice intake is the example: classify, extract, validate, approve, post. A workflow is cheaper, faster, testable and auditable. The model is used where judgement is needed (classification, extraction), and code owns the flow. Reach for an agent only when the path genuinely depends on what's discovered along the way.

Question 2Q: Single agent vs. multi-agent for a document-processing workflow: what does each side have going for it?Think it through, then reveal

A: For one agent (or a workflow): documents flow through the same steps, hand-offs lose context, multi-agent multiplies tokens, and one trace is far easier to debug. For several agents: documents are independent and can be processed in parallel, each document or section may overflow a single context, and specialists (tables, legal clauses, figures) benefit from focused instructions and tools. A common answer is a workflow that fans out per document to a focused worker, which is orchestrator-workers rather than free-form agents talking to each other.

Question 3Q: A router's classifier sometimes returns labels that aren't in its list. How should the code handle that?Think it through, then reveal

A: Normalize (case, whitespace), accept only known labels, and send everything else to a safe fallback. Log the misses and add them to the classifier's evaluation set. Consider structured output with an enum so the model can only produce valid labels.

Question 4Q: Why checkpoint after every step of a long workflow?Think it through, then reveal

A: So a crash costs one step, not the run. Resuming reuses completed work (no repeated model calls, no double-posted invoices when combined with idempotency keys) and avoids getting a different answer on the rerun.

Primary sources

The papers behind this lesson

Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022).

It established the reason-act-observe loop that the "agent loop" rung of the spectrum is built on.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Anthropic, Building effective agents (the named patterns): https://www.anthropic.com/engineering/building-effective-agents
  • LangGraph docs: https://langchain-ai.github.io/langgraph/
  • OpenAI Agents SDK: https://openai.github.io/openai-agents-python/
  • CrewAI docs: https://docs.crewai.com/
  • AutoGen: https://microsoft.github.io/autogen/
  • Temporal, durable execution: https://docs.temporal.io/

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.