At a glance
Key takeaways
- Success compounds: . 95% per step is 60% over ten steps.
- Plan first, execute step by step, replan from the point of failure, and keep finished work.
- Decompose into subtasks with explicit checks; retry only the failed step (checkpoints).
- Self-review shares the model's blind spots; tests, schemas and database checks don't.
- Retries help only for failures you detect, so the check matters more than the retry.
Level 2
How it works, from scratch
What follows measures the arithmetic, then builds the three remedies one mechanism at a time.
An agent that does one thing is easy. An agent that does twenty things in a row mostly fails, for a reason that is pure arithmetic. This lesson measures that arithmetic, then builds the three remedies: decompose the task into steps you can check, verify each step with something outside the model, and checkpoint so a failure costs one step instead of the whole run.
Chapter 1
Compounding error: why long tasks fail
Everyday picture A relay race with ten runners. Each runner drops the baton only 1 time in 20 during their leg. That sounds safe, but the team needs all ten legs to go cleanly, and it loses about 4 races in 10.
Tiny worked example Each step of an agent succeeds 95% of the time.
| steps | chance all succeed |
|---|---|
| 1 | 0.95 |
| 2 | 0.95 × 0.95 = 0.9025 |
| 10 | 0.95¹⁰ ≈ 0.599 |
| 20 | 0.95²⁰ ≈ 0.358 |
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning |
|---|---|
| chance a single step succeeds | |
| number of steps the task needs |
In words: multiply the per-step success rate by itself once per step.
On the example: and . A 95%-reliable step is a 60%-reliable 10-step task.
In Python:
p = 0.95
# ten steps that must all succeed
round(p ** 10, 2) # → 0.6
round(p ** 20, 2) # → 0.36
Figure 1 · Chart
At 95% per step a 10-step task succeeds only about 60% of the time, and every curve keeps sliding toward zero as steps are added
In code: end_to_end_success multiplies the per-step rate by itself once
per step, which is the whole formula above.
Why it matters The remedies all attack or : fewer steps
(higher-level tools, see primer.agents.tools), checks that catch failures,
retries of just the failed step, and human review at critical points.
Chapter 2
Plan-and-execute, with replanning
Everyday picture Cooking from a recipe. You read the whole recipe first (the plan), then do the steps. Halfway through you discover you're out of eggs. You don't start over, and you don't pretend you have eggs. You revise the rest of the recipe with what you've got, keeping everything you've already chopped.
Tiny worked example Goal "Reconcile Q3 invoices". The planner's first plan uses an old API. Here's the actual run:
planner -> {"steps": ["fetch_payments", "fetch_invoices_v1", "match", "draft_summary"]}
execute fetch_payments ok
execute fetch_invoices_v1 failed: invoices API v1 was retired on 2026-07-01; use fetch_invoices_v2
planner <- "Completed so far: fetch_payments. Step fetch_invoices_v1 failed: ...retired..."
planner -> {"steps": ["fetch_invoices_v2", "match", "draft_summary"]}
execute fetch_invoices_v2 ok
execute match ok
execute draft_summary ok (fetch_payments was NOT run again)
Figure 2 · Diagram
flowchart TD
G[Goal] --> P[Planner writes a plan]
P --> X[Execute next step]
X --> C{Step result<br/>meets expectations?}
C -->|yes| M{More steps?}
M -->|yes| X
M -->|no| D[Done]
C -->|no| R{Replans left?}
R -->|yes| F[Tell the planner:<br/>what's done + why it failed]
F --> P
R -->|no| S[Stop: report the failure]
The code plan_and_execute(planner, goal) asks for a JSON plan, runs
each action, records outputs in state, and on StepFailed asks for a
revised plan. Steps already in state are skipped.
In code: PlanRun is what plan_and_execute returns: every plan the
planner wrote, each executed step with "ok" or its failure, and the replan count.
Why it matters Deciding one step at a time (a pure ReAct loop, see
primer.agents.agent_loop) drifts on long tasks. An upfront plan keeps the
agent on track and lets a person see its intent before it acts. The risk is
following a plan that reality has invalidated, which is why replanning exists.
Chapter 3
Decomposition into verifiable subtasks
Everyday picture Moving house. "Move house" isn't something you can check off, but "every box labelled", "van booked for Saturday" and "keys handed back" are. Each has a clear definition of done.
Tiny worked example "Reconcile Q3 invoices" becomes five subtasks, each with a check:
| subtask | output | check (definition of done) |
|---|---|---|
| fetch_invoices | 4 invoices | at least one invoice returned |
| fetch_payments | 3 payments | at least one payment returned |
| match | 4 rows: invoiced vs paid | each invoice appears exactly once |
| list_mismatches | INV-102 (1200 vs 1100), INV-104 (450 vs 0) | none |
| draft_summary | "Q3 reconciliation: 4 invoices, 3 payments, 2 mismatches (INV-102 short by 100, INV-104 unpaid)." | mentions the mismatches |
When the payments API times out once, only fetch_payments runs again:
attempts are {fetch_invoices: 1, fetch_payments: 2, match: 1, ...}.
Figure 3 · Diagram
flowchart LR A[fetch_invoices] --> C[match] B[fetch_payments] --> C C --> D[list_mismatches] D --> E[draft_summary] B -. timeout .-> B
fetch_payments is a retry. Because every step's output is
kept (a checkpoint), the retry doesn't redo fetch_invoices.In code: Subtask pairs a step's work with its check (the definition of
done); run_subtasks runs them in order, retries only the step whose check
failed, and returns a RunReport of outputs and attempts. reconcile_q3
builds the five subtasks in the table.
Checkpoints: how much work a failure costs
Without checkpoints, a failure anywhere means starting again from step one. With them, you retry just the failed step. The expected number of step executions to finish an -step task:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning |
|---|---|
| chance a single step attempt succeeds (failures are noticed) | |
| number of steps |
In words: with checkpoints each step needs on average attempts, so steps need . Restarting from scratch needs successes in a row, and the expected wait for a run of successes grows roughly like .
On the example: , : with checkpoints step runs; restarting from scratch . At it's about 52.6 vs. 240.
In Python:
def with_checkpoints(n, p):
# each step needs 1/p attempts on average
return n / p
def restart_from_scratch(n, p):
return (1 - p ** n) / ((1 - p) * p ** n)
round(with_checkpoints(10, 0.95), 1), round(restart_from_scratch(10, 0.95), 1) # → (10.5, 13.4)
round(with_checkpoints(50, 0.95), 1), round(restart_from_scratch(50, 0.95)) # → (52.6, 240)
Figure 4 · Chart
On the log scale, restarting from scratch climbs as a straight line (exponential growth) while checkpoints stay close to one run per step
In code: expected_step_runs evaluates both formulas: with
checkpoints, the restart formula without.
Chapter 4
Reflection vs. external verification
Everyday picture Proofreading your own essay versus having someone run the numbers in it. You read what you meant to write, so your own blind spots stay blind. A calculator doesn't share them.
Tiny worked example The model writes a leap-year function:
def is_leap(year):
return year % 4 == 0 # forgets that 1900 was not a leap year
Reflection (asking a model to review it) replies "Looks correct: years
divisible by 4 are leap years." External verification (running it against
known answers) replies is_leap(1900) returned True, expected False. Fed
that failure, the model's second draft passes every case.
Figure 5 · Diagram
flowchart LR G[Model writes code] --> R1[Reflection:<br/>model reads its own code] R1 -->|same blind spot| A1[Approved, still wrong] G --> V[Verification:<br/>run tests / check schema / query the DB] V -->|1900 fails| F[Failure fed back] F --> G2[Second draft] G2 --> V2[Checks pass]
In code: self_review is the top path (a model reads the code and
approves or not); run_checks is the bottom path (it runs the code against
known answers); generate_until_checks_pass loops the bottom path, feeding
each failure back to the model until the checks pass.
With verification and retries, the per-step success rate rises:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning |
|---|---|
| chance one attempt succeeds | |
| chance the check detects a failed attempt | |
| retries allowed after a detected failure |
In words: you succeed on the first try, or fail and notice and succeed on the next try, and so on. Failures you don't notice get no retry.
On the example: , perfect check , one retry: per step, so ten steps succeed of the time instead of 60%. With a check that catches only half the failures (): .
In Python:
def p_step(p, d, r):
# p Σ_{i=0}^{r} ((1 - p) d)^i
return p * sum(((1 - p) * d) ** i for i in range(r + 1))
# a perfect check, one retry
round(p_step(0.95, 1, 1), 4) # → 0.9975
# ten such steps
round(p_step(0.95, 1, 1) ** 10, 3) # → 0.975
# a check that catches half the failures
round(p_step(0.95, 0.5, 1), 5) # → 0.97375
round(p_step(0.95, 0.5, 1) ** 10, 2) # → 0.77
Figure 6 · Chart
With the same 95% step, a ten-step task succeeds 97.5% of the time when a check catches every failure and allows one retry, 77% when the check catches half, and 60% with no checks
In code: step_success computes from , and ;
feed its result into end_to_end_success to get the ten-step curves.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Q: Each step of your agent is 95% reliable. How reliable is a 20-step task, and what do you do about it?Think it through, then reveal
A: . Cut the steps (higher-level tools), add external checks after key steps with a retry of just that step, checkpoint so failures don't restart the run, and put human review where a mistake is expensive. With perfect checks and one retry, per-step reliability becomes 99.75%, and 20 steps succeed about 95% of the time.
Question 2Q: Why isn't "ask the model to double-check its work" enough?Think it through, then reveal
A: The reviewer shares the author's blind spots, and self-critique can confidently reinforce a wrong answer. Checks outside the model (tests, schema validation, a query against the source of truth, a separate grader model with a rubric) produce concrete failures the model can act on.
Question 3Q: When is plan-and-execute better than deciding one step at a time?Think it through, then reveal
A: For long, multi-part tasks where staying on track matters and where showing the plan to a person before acting is valuable. Always pair it with replanning. A plan is a hypothesis, and the first failed step is evidence.
Question 4Q: What makes a subtask "good"?Think it through, then reveal
A: It's small, concrete and verifiable. It has a definition of done that code can check, and its output is saved so a later failure doesn't redo it.
Primary sources
The papers behind this lesson
It showed that asking a model to first devise a plan and then carry it out step by step reduces missed steps compared with reasoning straight through.
The paper ↗Agents that turn feedback from the environment (failed tests, wrong answers) into written lessons and retry improve markedly, and the gains come from external signals.
Read the annotated companion →The paper ↗Researcher's shelf
Further reading
- Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
- Lilian Weng, LLM Powered Autonomous Agents (planning, reflection): https://lilianweng.github.io/posts/2023-06-23-agent/
- Temporal, durable execution: https://docs.temporal.io/
- LangGraph docs (persistence and checkpoints): https://langchain-ai.github.io/langgraph/
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.