rumblr Work in progressWIP

● The AI Primer · Lesson 53 · Part 2: building systems people rely on

Why the hard ones fail

a catalogue of agent failures, and the fix for each

This lesson covers The common failure modes, and the fix for each

Members · open during launch 17 min6 figures and diagrams
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Long tasks fail multiplicatively: at 95% per step, 10 steps succeed 60% of the time. Shorten chains and verify at key steps.
  2. Retry transient failures with capped exponential backoff and jitter; never retry invalid requests; make writes idempotent first.
  3. Put circuit breakers around dependencies so outages fail fast instead of cascading.
  4. Detect loops (same call, same arguments) and enforce step and token budgets.
  5. Most failures are system failures (retrieval, tools, data, evals, adoption), not model failures, and each has a known fix.

Level 2

How it works, from scratch

Pilots learn from a catalogue of accidents: each entry names what went wrong, how it showed up in the cockpit, and the checklist item that now prevents it. This lesson is that catalogue for agents: ten failures that keep recurring in production systems, each with its symptom, its fix, and a pointer to the lesson that builds the fix. Three of them get their fix built right here: compounding error (arithmetic), brittle integrations (retries with backoff, and circuit breakers) and runaway loops (loop detection).

Failure Symptom Fix Lesson
Compounding error Long tasks fail far more often than short ones Fewer steps, checks after key steps, checkpoints primer.agents.planning
Bad retrieval Confident, wrong answers Hybrid search, reranking, retrieval evals primer.agents.rag
Ambiguous tools Wrong tool, or bad arguments Precise descriptions, fewer tools, validation primer.agents.tools
Context rot Quality drops as a session grows Summarize, trim, restart with a state handoff primer.agents.context
Loops and runaway The same call repeated; cost spikes Step budgets, loop detection, stop conditions primer.agents.agent_loop
No evals Regressions ship silently Golden sets in CI, online monitoring primer.agents.evals
Prompt injection The agent obeys instructions found in data Untrusted-content boundaries, privilege separation primer.agents.guardrails
Messy enterprise data Garbled tables, missing permissions Invest in parsing; permission-aware retrieval primer.agents.rag
Brittle integrations Timeouts, expired credentials, API changes Retries with backoff, circuit breakers, contract tests primer.agents.failures
No adoption It works, and nobody uses it Build with users, show sources, easy human handoff primer.agents.deployment

In code: FAILURES is this table as data, one Failure record per row.

Chapter 1

Compounding error

Everyday picture A relay race where each baton pass succeeds 95% of the time. One pass is nearly safe; ten passes in a row drop the baton more often than you'd think.

Worked example If each step of an agent's task succeeds 95% of the time, independently:

Steps Chance every step succeeds
1 0.95
2 0.95 × 0.95 = 0.9025
10 0.95¹⁰ ≈ 0.599
20 0.95²⁰ ≈ 0.358
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
chance a single step succeeds 0.95
number of steps that must all succeed 10
multiplied by itself times 0.599

In words: the chance that every step succeeds is the per-step chance multiplied together once per step.

On the worked example: 0.95 multiplied by itself 10 times is 0.599, so a ten-step task at 95% per step fails four times in ten.

Level 3: in Python
p, n = 0.95, 10
# p multiplied by itself n times
round(p ** n, 3)  # → 0.599
# fails about four times in ten
round(1 - p ** n, 1)  # → 0.4

Figure 1 · Diagram

Reading it: every arrow is another chance to drop the baton, so success shrinks with each step. The dotted loop is the remedy: check a step's result where it happens and retry only that step, so an error doesn't travel down the chain. Fewer steps, verification after key steps, checkpoints to resume from the middle, and human review at critical points all attack the same exponent.

Figure 2 · Chart

0 5 10 15 20 25 30 steps that must all succeed 0.0 0.2 0.4 0.6 0.8 1.0 chance the whole task succeeds Compounding error 60% 36% 99% per step 95% per step 90% per step

Twenty steps succeed 82% of the time at 99% per step but only 36% at 95% and 12% at 90%

Reading it: the x-axis is the number of steps; each curve is a per-step reliability. At 99% per step a 20-step task still succeeds 82% of the time; at 95% it's 36%; at 90%, 12%. Small gains in per-step reliability are worth far more than they look, and demos with three steps say little about tasks with twenty.

In code: chain_success is ; max_steps_for runs it backwards, returning the most steps you can chain and still reach a target success rate.

Chapter 2

Brittle integrations: retries with backoff

Everyday picture Calling a busy phone line: you don't redial every second. You wait a little, then longer, then longer still, and if everyone redials on the same schedule the line stays jammed, so you add a random pause (jitter).

Worked example A service times out twice and then answers. With a base delay of 0.1 s the waits are 0.1 s, then 0.2 s, and the third attempt succeeds. With a cap of 10 s, the waits never exceed 10 s however many attempts there are. A request that's invalid (a ValueError here) fails the same way every time, so it's raised immediately: retrying only adds load.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
wait before retry number = 0.1 s, = 0.2 s
how many retries have already happened 0, 1
base delay 0.1 s
doubles with every retry 1, 2, 4, …
the cap on any single wait 10 s

In words: each wait doubles the previous one, starting from the base delay, but never exceeds the cap.

On the worked example: = min(10, 0.1 × 1) = 0.1 s and = min(10, 0.1 × 2) = 0.2 s.

Level 3: in Python
b, D_max = 0.1, 10
# d_0 and d_1
[min(D_max, b * 2 ** k) for k in range(2)]  # → [0.1, 0.2]

Figure 3 · Diagram

Reading it: time runs downward. Each failure is followed by a longer pause, which gives an overloaded service room to recover instead of being hammered. Writes must be idempotent (safe to repeat: sending the same request twice has the same effect as once, usually via an idempotency key) before you retry them, or a retried "create order" makes two orders.

Figure 4 · Chart

1 2 3 4 5 6 7 retry number 0 2 4 6 8 10 seconds to wait Backoff: base 0.5 s, doubling, capped at 10 s capped exponential (no jitter) five clients with full jitter

The delay doubles each attempt until it hits the 10-second cap, and jitter scatters five clients' retries across the whole range below it

Reading it: the x-axis is the attempt number; the black line is the capped exponential delay. The dots are five clients using full jitter (each waits a random time between zero and the capped delay): instead of all retrying at the same moment, their retries spread out, which is what lets a recovering service actually recover.

In code: backoff_delays lists the capped waits ; retry_with_backoff calls a function, retries only the exceptions in RETRYABLE with those waits (optionally with full jitter), and raises anything else at once.

Chapter 3

Brittle integrations: circuit breakers

Everyday picture The breaker in your home's fuse box. When a circuit keeps shorting, it trips and cuts the power, instead of letting the wire overheat. After a while you flip it back on to test; if it trips again, it stays off.

A circuit breaker wraps calls to a dependency. After several consecutive failures it opens: calls fail immediately without touching the struggling service (fail fast), so the agent can fall back (use a cache, tell the user, hand off to a person) instead of hanging on timeouts. After a cool-down it goes half-open and lets one trial call through: success closes it, failure opens it again.

Worked example Threshold 3, cool-down 30 s. Three timeouts in a row: open. A fourth call at 5 s fails instantly with CircuitOpen and the service isn't called at all. At 31 s, one trial call goes through; it succeeds, so the breaker closes.

Figure 5 · Diagram

Reading it: in the normal state (Closed) calls flow through and failures are counted. Open is the protective state: nothing reaches the dependency. Half-open is a single, cautious test. The breaker turns a slow, cascading failure (every request waiting 30 s on a dead service) into a fast, contained one.

Figure 6 · Chart

0 20 40 60 80 100 seconds Circuit breaker (threshold 3, cool-down 15 s) during an outage service down ok failed failed fast (breaker open)

Through a 50-second outage the failing service is called only 5 times, while 21 requests fail fast at the open breaker instead

Reading it: the shaded band is when the upstream service is down; each marker is one request. Before the outage, calls succeed (green). The first three failures (red) open the breaker; after that, requests fail fast (grey) without reaching the service, apart from one trial call per cool-down. Once the service is back, the next trial succeeds and traffic resumes. The service received a handful of calls during the outage instead of all of them.

In code: CircuitBreaker.call is the state diagram: it raises CircuitOpen while open, lets one trial call through after the cool-down, and closes on success. simulate_outage replays the outage in the figure.

Contract tests (automated checks that an external API still accepts and returns what your tool expects) catch the third kind of brittleness, upstream API changes, before users do.

Chapter 4

Loops and runaway

Everyday picture A satnav that keeps rerouting you around the same block. The fix isn't a better map; it's noticing that you've passed the same corner three times.

Worked example search("vpn"), search("vpn"), search("vpn"): the same tool with the same arguments three times in a row. Nothing new can come back, so the agent is looping. search("vpn"), search("vpn error"), search("ERR-4012") is the same tool refining its query: progress. Combine loop detection with hard step and token budgets (primer.agents.cost.TaskBudget) so a confused agent stops and hands off instead of burning money.

In code: is_looping reports whether the last few calls were the same tool with identical arguments.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Your agent works in demos and fails half the time in production. Where do you look first?Think it through, then reveal

At the length of real tasks versus demo tasks (compounding error), and at traces for the first failing step: retrieval misses, wrong tools or arguments, tool errors from integrations, context growth. Then fix per-step reliability where it's lowest, add verification after key steps, and put the failing cases into the eval set.

Question 2Why add jitter to retries?Think it through, then reveal

Without it, every client that failed at the same moment retries at the same moment, producing synchronized waves of load that keep an overloaded service down. Random delays spread the retries out.

Question 3When should you not retry?Think it through, then reveal

When the error can't go away on its own: invalid input, permission denied, not found. And never retry a non-idempotent write without an idempotency key, or you'll duplicate its effect.

Question 4What's the difference between a retry and a circuit breaker?Think it through, then reveal

A retry handles a blip for one request. A breaker handles an outage across many requests: it stops sending traffic to a failing dependency, fails fast so callers can fall back, and probes for recovery.

Primary sources

The papers behind this lesson

Researcher's shelf

Further reading

  • Martin Fowler, CircuitBreaker: https://martinfowler.com/bliki/CircuitBreaker.html
  • AWS Builders' Library, Timeouts, retries, and backoff with jitter: https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/
  • Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.