At a glance
Key takeaways
- Long tasks fail multiplicatively: at 95% per step, 10 steps succeed 60% of the time. Shorten chains and verify at key steps.
- Retry transient failures with capped exponential backoff and jitter; never retry invalid requests; make writes idempotent first.
- Put circuit breakers around dependencies so outages fail fast instead of cascading.
- Detect loops (same call, same arguments) and enforce step and token budgets.
- Most failures are system failures (retrieval, tools, data, evals, adoption), not model failures, and each has a known fix.
Level 2
How it works, from scratch
Pilots learn from a catalogue of accidents: each entry names what went wrong, how it showed up in the cockpit, and the checklist item that now prevents it. This lesson is that catalogue for agents: ten failures that keep recurring in production systems, each with its symptom, its fix, and a pointer to the lesson that builds the fix. Three of them get their fix built right here: compounding error (arithmetic), brittle integrations (retries with backoff, and circuit breakers) and runaway loops (loop detection).
| Failure | Symptom | Fix | Lesson |
|---|---|---|---|
| Compounding error | Long tasks fail far more often than short ones | Fewer steps, checks after key steps, checkpoints | primer.agents.planning |
| Bad retrieval | Confident, wrong answers | Hybrid search, reranking, retrieval evals | primer.agents.rag |
| Ambiguous tools | Wrong tool, or bad arguments | Precise descriptions, fewer tools, validation | primer.agents.tools |
| Context rot | Quality drops as a session grows | Summarize, trim, restart with a state handoff | primer.agents.context |
| Loops and runaway | The same call repeated; cost spikes | Step budgets, loop detection, stop conditions | primer.agents.agent_loop |
| No evals | Regressions ship silently | Golden sets in CI, online monitoring | primer.agents.evals |
| Prompt injection | The agent obeys instructions found in data | Untrusted-content boundaries, privilege separation | primer.agents.guardrails |
| Messy enterprise data | Garbled tables, missing permissions | Invest in parsing; permission-aware retrieval | primer.agents.rag |
| Brittle integrations | Timeouts, expired credentials, API changes | Retries with backoff, circuit breakers, contract tests | primer.agents.failures |
| No adoption | It works, and nobody uses it | Build with users, show sources, easy human handoff | primer.agents.deployment |
In code: FAILURES is this table as data, one Failure record per row.
Chapter 1
Compounding error
Everyday picture A relay race where each baton pass succeeds 95% of the time. One pass is nearly safe; ten passes in a row drop the baton more often than you'd think.
Worked example If each step of an agent's task succeeds 95% of the time, independently:
| Steps | Chance every step succeeds |
|---|---|
| 1 | 0.95 |
| 2 | 0.95 × 0.95 = 0.9025 |
| 10 | 0.95¹⁰ ≈ 0.599 |
| 20 | 0.95²⁰ ≈ 0.358 |
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| chance a single step succeeds | 0.95 | |
| number of steps that must all succeed | 10 | |
| multiplied by itself times | 0.599 |
In words: the chance that every step succeeds is the per-step chance multiplied together once per step.
On the worked example: 0.95 multiplied by itself 10 times is 0.599, so a ten-step task at 95% per step fails four times in ten.
Level 3: in Python
p, n = 0.95, 10
# p multiplied by itself n times
round(p ** n, 3) # → 0.599
# fails about four times in ten
round(1 - p ** n, 1) # → 0.4
Figure 1 · Diagram
flowchart LR S1[Step 1<br/>95%] --> S2[Step 2<br/>95%] --> S3[...] --> S10[Step 10<br/>95%] --> D[Done:<br/>60% of the time] S2 -.->|verify, retry<br/>just this step| S2
Figure 2 · Chart
Twenty steps succeed 82% of the time at 99% per step but only 36% at 95% and 12% at 90%
In code: chain_success is ; max_steps_for runs it backwards,
returning the most steps you can chain and still reach a target success rate.
Chapter 2
Brittle integrations: retries with backoff
Everyday picture Calling a busy phone line: you don't redial every second. You wait a little, then longer, then longer still, and if everyone redials on the same schedule the line stays jammed, so you add a random pause (jitter).
Worked example A service times out twice and then answers. With a base
delay of 0.1 s the waits are 0.1 s, then 0.2 s, and the third attempt
succeeds. With a cap of 10 s, the waits never exceed 10 s however many
attempts there are. A request that's invalid (a ValueError here) fails
the same way every time, so it's raised immediately: retrying only adds
load.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| wait before retry number | = 0.1 s, = 0.2 s | |
| how many retries have already happened | 0, 1 | |
| base delay | 0.1 s | |
| doubles with every retry | 1, 2, 4, … | |
| the cap on any single wait | 10 s |
In words: each wait doubles the previous one, starting from the base delay, but never exceeds the cap.
On the worked example: = min(10, 0.1 × 1) = 0.1 s and = min(10, 0.1 × 2) = 0.2 s.
Level 3: in Python
b, D_max = 0.1, 10
# d_0 and d_1
[min(D_max, b * 2 ** k) for k in range(2)] # → [0.1, 0.2]
Figure 3 · Diagram
sequenceDiagram participant A as Agent tool participant S as Upstream API A->>S: request S--xA: timeout Note over A: wait 0.1 s A->>S: retry 1 S--xA: timeout Note over A: wait 0.2 s A->>S: retry 2 S-->>A: 200 OK
Figure 4 · Chart
The delay doubles each attempt until it hits the 10-second cap, and jitter scatters five clients' retries across the whole range below it
In code: backoff_delays lists the capped waits ;
retry_with_backoff calls a function, retries only the exceptions in
RETRYABLE with those waits (optionally with full jitter), and raises
anything else at once.
Chapter 3
Brittle integrations: circuit breakers
Everyday picture The breaker in your home's fuse box. When a circuit keeps shorting, it trips and cuts the power, instead of letting the wire overheat. After a while you flip it back on to test; if it trips again, it stays off.
A circuit breaker wraps calls to a dependency. After several consecutive failures it opens: calls fail immediately without touching the struggling service (fail fast), so the agent can fall back (use a cache, tell the user, hand off to a person) instead of hanging on timeouts. After a cool-down it goes half-open and lets one trial call through: success closes it, failure opens it again.
Worked example Threshold 3, cool-down 30 s. Three timeouts in a row:
open. A fourth call at 5 s fails instantly with CircuitOpen and the
service isn't called at all. At 31 s, one trial call goes through; it
succeeds, so the breaker closes.
Figure 5 · Diagram
stateDiagram-v2 [*] --> Closed Closed --> Closed: success (reset the failure count) Closed --> Open: 3 failures in a row Open --> Open: calls rejected instantly Open --> HalfOpen: cool-down elapsed HalfOpen --> Closed: trial call succeeds HalfOpen --> Open: trial call fails
Figure 6 · Chart
Through a 50-second outage the failing service is called only 5 times, while 21 requests fail fast at the open breaker instead
In code: CircuitBreaker.call is the state diagram: it raises
CircuitOpen while open, lets one trial call through after the cool-down,
and closes on success. simulate_outage replays the outage in the figure.
Contract tests (automated checks that an external API still accepts and returns what your tool expects) catch the third kind of brittleness, upstream API changes, before users do.
Chapter 4
Loops and runaway
Everyday picture A satnav that keeps rerouting you around the same block. The fix isn't a better map; it's noticing that you've passed the same corner three times.
Worked example search("vpn"), search("vpn"), search("vpn"): the
same tool with the same arguments three times in a row. Nothing new can
come back, so the agent is looping. search("vpn"), search("vpn error"),
search("ERR-4012") is the same tool refining its query: progress.
Combine loop detection with hard step and token budgets
(primer.agents.cost.TaskBudget) so a confused agent stops and hands off
instead of burning money.
In code: is_looping reports whether the last few calls were the same
tool with identical arguments.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Your agent works in demos and fails half the time in production. Where do you look first?Think it through, then reveal
At the length of real tasks versus demo tasks (compounding error), and at traces for the first failing step: retrieval misses, wrong tools or arguments, tool errors from integrations, context growth. Then fix per-step reliability where it's lowest, add verification after key steps, and put the failing cases into the eval set.
Question 2Why add jitter to retries?Think it through, then reveal
Without it, every client that failed at the same moment retries at the same moment, producing synchronized waves of load that keep an overloaded service down. Random delays spread the retries out.
Question 3When should you not retry?Think it through, then reveal
When the error can't go away on its own: invalid input, permission denied, not found. And never retry a non-idempotent write without an idempotency key, or you'll duplicate its effect.
Question 4What's the difference between a retry and a circuit breaker?Think it through, then reveal
A retry handles a blip for one request. A breaker handles an outage across many requests: it stops sending traffic to a failing dependency, fails fast so callers can fall back, and probes for recovery.
Primary sources
The papers behind this lesson
Researcher's shelf
Further reading
- Martin Fowler, CircuitBreaker: https://martinfowler.com/bliki/CircuitBreaker.html
- AWS Builders' Library, Timeouts, retries, and backoff with jitter: https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/
- Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.