rumblr Work in progressWIP

● The AI Primer · Lesson 47 · Part 2: building systems people rely on

Multi-step planning

plans, checks, retries and why long tasks fail

This lesson covers Plan-and-execute, decomposition, reflection, compounding error

Members · open during launch 19 min6 figures and diagrams
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Success compounds: . 95% per step is 60% over ten steps.
  2. Plan first, execute step by step, replan from the point of failure, and keep finished work.
  3. Decompose into subtasks with explicit checks; retry only the failed step (checkpoints).
  4. Self-review shares the model's blind spots; tests, schemas and database checks don't.
  5. Retries help only for failures you detect, so the check matters more than the retry.

Level 2

How it works, from scratch

What follows measures the arithmetic, then builds the three remedies one mechanism at a time.

An agent that does one thing is easy. An agent that does twenty things in a row mostly fails, for a reason that is pure arithmetic. This lesson measures that arithmetic, then builds the three remedies: decompose the task into steps you can check, verify each step with something outside the model, and checkpoint so a failure costs one step instead of the whole run.

Chapter 1

Compounding error: why long tasks fail

Everyday picture A relay race with ten runners. Each runner drops the baton only 1 time in 20 during their leg. That sounds safe, but the team needs all ten legs to go cleanly, and it loses about 4 races in 10.

Tiny worked example Each step of an agent succeeds 95% of the time.

steps chance all succeed
1 0.95
2 0.95 × 0.95 = 0.9025
10 0.95¹⁰ ≈ 0.599
20 0.95²⁰ ≈ 0.358
Level 3: the formula and its symbols

Symbols

Symbol Meaning
chance a single step succeeds
number of steps the task needs

In words: multiply the per-step success rate by itself once per step.

On the example: and . A 95%-reliable step is a 60%-reliable 10-step task.

In Python:

p = 0.95
# ten steps that must all succeed
round(p ** 10, 2)  # → 0.6
round(p ** 20, 2)  # → 0.36

Figure 1 · Drawn from the lesson's code

0 5 10 15 20 25 30 steps in the task (n) 0.0 0.2 0.4 0.6 0.8 1.0 P(task succeeds) = p^n Compounding error: reliable steps, unreliable tasks 90% per step 95% per step 99% per step

At 95% per step a 10-step task succeeds only about 60% of the time, and every curve keeps sliding toward zero as steps are added

Reading it: the x-axis is the number of steps in the task, and the y-axis is the chance the whole task succeeds. Every curve starts near the top and slides down, and the lower the per-step rate, the faster it falls. Find 10 steps on the x-axis and read up to the 95% curve: about 0.6. That's why a demo of a three-step task looks great and the same agent on a real twenty-step job disappoints.

In code: end_to_end_success multiplies the per-step rate by itself once per step, which is the whole formula above.

Why it matters The remedies all attack or : fewer steps (higher-level tools, see primer.agents.tools), checks that catch failures, retries of just the failed step, and human review at critical points.

Chapter 2

Plan-and-execute, with replanning

Everyday picture Cooking from a recipe. You read the whole recipe first (the plan), then do the steps. Halfway through you discover you're out of eggs. You don't start over, and you don't pretend you have eggs. You revise the rest of the recipe with what you've got, keeping everything you've already chopped.

Tiny worked example Goal "Reconcile Q3 invoices". The planner's first plan uses an old API. Here's the actual run:

planner  -> {"steps": ["fetch_payments", "fetch_invoices_v1", "match", "draft_summary"]}
execute  fetch_payments      ok
execute  fetch_invoices_v1   failed: invoices API v1 was retired on 2026-07-01; use fetch_invoices_v2
planner  <- "Completed so far: fetch_payments. Step fetch_invoices_v1 failed: ...retired..."
planner  -> {"steps": ["fetch_invoices_v2", "match", "draft_summary"]}
execute  fetch_invoices_v2   ok
execute  match               ok
execute  draft_summary       ok     (fetch_payments was NOT run again)

Figure 2 · Diagram

Reading it: the inner loop (execute → check → next) is the plan being followed. The outer loop back to Planner only happens on a failed check, and it carries two facts: what's already done (so work is kept) and why the step failed (so the new plan can route around it). The Replans left? diamond stops a planner that can't adapt from looping forever.

The code plan_and_execute(planner, goal) asks for a JSON plan, runs each action, records outputs in state, and on StepFailed asks for a revised plan. Steps already in state are skipped.

In code: PlanRun is what plan_and_execute returns: every plan the planner wrote, each executed step with "ok" or its failure, and the replan count.

Why it matters Deciding one step at a time (a pure ReAct loop, see primer.agents.agent_loop) drifts on long tasks. An upfront plan keeps the agent on track and lets a person see its intent before it acts. The risk is following a plan that reality has invalidated, which is why replanning exists.

Chapter 3

Decomposition into verifiable subtasks

Everyday picture Moving house. "Move house" isn't something you can check off, but "every box labelled", "van booked for Saturday" and "keys handed back" are. Each has a clear definition of done.

Tiny worked example "Reconcile Q3 invoices" becomes five subtasks, each with a check:

subtask output check (definition of done)
fetch_invoices 4 invoices at least one invoice returned
fetch_payments 3 payments at least one payment returned
match 4 rows: invoiced vs paid each invoice appears exactly once
list_mismatches INV-102 (1200 vs 1100), INV-104 (450 vs 0) none
draft_summary "Q3 reconciliation: 4 invoices, 3 payments, 2 mismatches (INV-102 short by 100, INV-104 unpaid)." mentions the mismatches

When the payments API times out once, only fetch_payments runs again: attempts are {fetch_invoices: 1, fetch_payments: 2, match: 1, ...}.

Figure 3 · Diagram

Reading it: arrows show which outputs feed which step. The dotted self-loop on fetch_payments is a retry. Because every step's output is kept (a checkpoint), the retry doesn't redo fetch_invoices.

In code: Subtask pairs a step's work with its check (the definition of done); run_subtasks runs them in order, retries only the step whose check failed, and returns a RunReport of outputs and attempts. reconcile_q3 builds the five subtasks in the table.

Checkpoints: how much work a failure costs

Without checkpoints, a failure anywhere means starting again from step one. With them, you retry just the failed step. The expected number of step executions to finish an -step task:

Level 3: the formula and its symbols

Symbols

Symbol Meaning
chance a single step attempt succeeds (failures are noticed)
number of steps

In words: with checkpoints each step needs on average attempts, so steps need . Restarting from scratch needs successes in a row, and the expected wait for a run of successes grows roughly like .

On the example: , : with checkpoints step runs; restarting from scratch . At it's about 52.6 vs. 240.

In Python:

def with_checkpoints(n, p):
    # each step needs 1/p attempts on average
    return n / p
def restart_from_scratch(n, p):
    return (1 - p ** n) / ((1 - p) * p ** n)
round(with_checkpoints(10, 0.95), 1), round(restart_from_scratch(10, 0.95), 1)  # → (10.5, 13.4)
round(with_checkpoints(50, 0.95), 1), round(restart_from_scratch(50, 0.95))  # → (52.6, 240)

Figure 4 · Drawn from the lesson's code

0 10 20 30 40 50 60 steps in the task (n) 1 0 0 1 0 1 1 0 2 expected step executions (log scale) What a failure costs (p = 0.95 per step) checkpoints: retry the failed step (n/p) no checkpoints: restart from step 1

On the log scale, restarting from scratch climbs as a straight line (exponential growth) while checkpoints stay close to one run per step

Reading it: the x-axis is task length, and the y-axis (log scale) is how many step executions you should expect to pay for. On a log scale each gridline is ten times the one below, so exponential growth draws a straight line. The restart line is that straight line: every extra step multiplies its cost by about . The checkpoint line (, just proportional to ) bends over and flattens, because on a log scale going from 10 to 20 steps rises no more than going from 5 to 10. The widening gap between them is the cost of having no checkpoints, and at 50 steps it's already about 240 step runs against 53. Long tasks without checkpoints aren't just unreliable; they're expensive. Durable execution is the engineering name for this: a workflow engine that saves each step's result so a crashed run resumes where it stopped.

In code: expected_step_runs evaluates both formulas: with checkpoints, the restart formula without.

Chapter 4

Reflection vs. external verification

Everyday picture Proofreading your own essay versus having someone run the numbers in it. You read what you meant to write, so your own blind spots stay blind. A calculator doesn't share them.

Tiny worked example The model writes a leap-year function:

def is_leap(year):
    return year % 4 == 0        # forgets that 1900 was not a leap year

Reflection (asking a model to review it) replies "Looks correct: years divisible by 4 are leap years." External verification (running it against known answers) replies is_leap(1900) returned True, expected False. Fed that failure, the model's second draft passes every case.

Figure 5 · Diagram

Reading it: two paths leave the same first draft. The top path stays inside the model and ends at "approved, still wrong". The bottom path goes through something outside the model (tests, a schema validator, a database query), and that something produces a concrete, checkable failure the model can fix.

In code: self_review is the top path (a model reads the code and approves or not); run_checks is the bottom path (it runs the code against known answers); generate_until_checks_pass loops the bottom path, feeding each failure back to the model until the checks pass.

With verification and retries, the per-step success rate rises:

Level 3: the formula and its symbols

Symbols

Symbol Meaning
chance one attempt succeeds
chance the check detects a failed attempt
retries allowed after a detected failure

In words: you succeed on the first try, or fail and notice and succeed on the next try, and so on. Failures you don't notice get no retry.

On the example: , perfect check , one retry: per step, so ten steps succeed of the time instead of 60%. With a check that catches only half the failures (): .

In Python:

def p_step(p, d, r):
    # p Σ_{i=0}^{r} ((1 - p) d)^i
    return p * sum(((1 - p) * d) ** i for i in range(r + 1))
# a perfect check, one retry
round(p_step(0.95, 1, 1), 4)  # → 0.9975
# ten such steps
round(p_step(0.95, 1, 1) ** 10, 3)  # → 0.975
# a check that catches half the failures
round(p_step(0.95, 0.5, 1), 5)  # → 0.97375
round(p_step(0.95, 0.5, 1) ** 10, 2)  # → 0.77

Figure 6 · Drawn from the lesson's code

0 5 10 15 20 25 30 steps in the task (n) 0.0 0.2 0.4 0.6 0.8 1.0 P(task succeeds) Same 95% step, different verification no checks (step = 0.9500) check catches half, 1 retry (step = 0.9737) check catches all, 1 retry (step = 0.9975) check catches all, 2 retries (step = 0.9999)

With the same 95% step, a ten-step task succeeds 97.5% of the time when a check catches every failure and allows one retry, 77% when the check catches half, and 60% with no checks

Reading it: all curves use the same 95%-reliable step. The bottom curve has no checks. The middle ones add a retry after a failure is caught, with a check that catches half or all failures. The top curve allows two retries. The gap between "half" and "all" is the lesson: retries are only as good as the check that triggers them. Invest in the check.

In code: step_success computes from , and ; feed its result into end_to_end_success to get the ten-step curves.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: Each step of your agent is 95% reliable. How reliable is a 20-step task, and what do you do about it?Think it through, then reveal

A: . Cut the steps (higher-level tools), add external checks after key steps with a retry of just that step, checkpoint so failures don't restart the run, and put human review where a mistake is expensive. With perfect checks and one retry, per-step reliability becomes 99.75%, and 20 steps succeed about 95% of the time.

Question 2Q: Why isn't "ask the model to double-check its work" enough?Think it through, then reveal

A: The reviewer shares the author's blind spots, and self-critique can confidently reinforce a wrong answer. Checks outside the model (tests, schema validation, a query against the source of truth, a separate grader model with a rubric) produce concrete failures the model can act on.

Question 3Q: When is plan-and-execute better than deciding one step at a time?Think it through, then reveal

A: For long, multi-part tasks where staying on track matters and where showing the plan to a person before acting is valuable. Always pair it with replanning. A plan is a hypothesis, and the first failed step is evidence.

Question 4Q: What makes a subtask "good"?Think it through, then reveal

A: It's small, concrete and verifiable. It has a definition of done that code can check, and its output is saved so a later failure doesn't redo it.

Primary sources

The papers behind this lesson

Wang et al., Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models (2023).

It showed that asking a model to first devise a plan and then carry it out step by step reduces missed steps compared with reasoning straight through.

The paper ↗
Shinn et al., Reflexion: Language Agents with Verbal Reinforcement Learning (2023).

Agents that turn feedback from the environment (failed tests, wrong answers) into written lessons and retry improve markedly, and the gains come from external signals.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
  • Lilian Weng, LLM Powered Autonomous Agents (planning, reflection): https://lilianweng.github.io/posts/2023-06-23-agent/
  • Temporal, durable execution: https://docs.temporal.io/
  • LangGraph docs (persistence and checkpoints): https://langchain-ai.github.io/langgraph/

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.