rumblr Work in progressWIP

● The AI Primer · Lesson 40 · Part 2: building systems people rely on

The agent loop, production-grade

You'll be able to explain A production agent loop: budgets, loop detection, recovery

Members · open during launch 19 min7 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. The agent loop is: call model, run the tools it asks for, append results, repeat until end_turn.
  2. The model never executes anything. Your code validates and runs every tool call.
  3. Production loops need step limits, token/cost budgets, loop detection, error-as-result handling, and a human hand-off.
  4. Parallel tool calls: run them concurrently, return all results in a single user message.
  5. Input cost grows roughly with the square of the number of steps, because every call re-sends the history.

Level 1

The practitioner's guide

In one sentence

The agent loop is the piece of code that calls the model, runs the tools it asks for, sends the results back and repeats until the model says it is done or a limit you set says it is done for it.

When you need it

You need a loop whenever the model has to take more than one step whose order it decides: look something up, then act on what it found, then check. One call with one tool call you run once is not a loop, and neither is a fixed pipeline where your code decides the steps (primer.agents.orchestration). You need the production loop, rather than the ten-line version, the moment real money and real side effects are attached, because the ten-line version has no way to stop. In this lesson's happy path, one question ("How many PTO days does alice have left, and what rolls over?") takes 2 model calls and 2 tool calls run in parallel, and stops because the model said end_turn. In the same lesson, a model that keeps searching stops only because a token budget tripped: 11 calls in, the running total passes 12,000 tokens (12,990) and the run ends with an estimated cost of $0.07. The tell that you need the controls: you cannot say, for every run that ended, why it ended.

Your options

Six ways to run the loop, from the least machinery to the most:

Option What it does What it guarantees What it costs Where it lives
No loop One model call; if it asks for a tool, run it once and call once more Bounded cost: at most two calls Nothing that needs a second decision gets done Your code
The SDK's tool runner Drives the request, run, reply cycle until the model returns no tool call or max_iterations is reached The plumbing done right: results paired by id, one message per turn, type-checked inputs Per-turn controls are limited; Claude's docs point you to the manual loop for approval, custom logging or conditional execution The provider's SDK (seven languages for Claude)
The ten-line loop Call, run tools, append results, repeat while stop_reason is tool_use Full control of every turn Every production failure is yours: it can run forever, repeat itself, crash on a tool error or execute a half-written call Your code
The production loop (this lesson) The ten-line loop plus a step limit, token and dollar budgets, loop detection, errors returned as results, a hand-off tool and parallel tool execution A named stop cause for every run, one of seven A config to tune (steps, tokens, dollars, repeat threshold) and a test per control Your code
A framework's loop The loop with tracing, persistence and hand-offs built in, and its own limit (max_turns, a recursion limit) Standard patterns and observability without writing them A dependency, its assumptions and its defaults (LangGraph's recursion limit defaults to 1,000 steps) LangGraph, the OpenAI Agents SDK, the Claude Agent SDK
A hosted loop The provider runs the loop and a sandbox for the tools; you send messages and tool results No loop code, no state files, per-session containers Session runtime on top of tokens (Claude Managed Agents lists $0.08 per session-hour) and less say over each turn The provider's servers

How to choose

Start from what can go wrong and who pays for it.

  • A prototype, a notebook, a one-off script: the tool runner. It gets the message pairing right, which is the part people get wrong first.
  • Production with side effects (writes, emails, refunds): own the loop, or use a framework whose per-turn hooks you have read. You need an approval gate, a budget in dollars, and a hand-off to a person, in exactly your shape.
  • Many independent tool calls per turn (three lookups): make sure whatever you use runs them concurrently and returns all results in one message. The lesson's latency figure shows serial time climbing with every tool while concurrent time flattens at the slowest call.
  • Long-running or scheduled agents you would rather not host: a hosted loop, if its controls cover your approval and budget needs.
  • Whatever you pick, set every limit before the first real run: steps, tokens, dollars and a repeat threshold. This lesson's defaults are 10 steps, 50,000 tokens, and a stop when the same tool is asked for with the same arguments 3 times.

What it costs

Input tokens grow with the square of the number of steps, because each call re-sends the whole history: with 500 new tokens a step, 10 steps send 27,500 input tokens rather than the 5,000 a flat per-step cost suggests, and 20 steps send 105,000, almost four times ten steps' total. At Claude Opus 5's list prices ($5 per million input tokens and $25 per million output, the defaults in this lesson's AgentConfig), the runaway run above cost about 7 cents before its 12,000-token budget stopped it. Latency is one round trip per model call plus the slowest tool in each turn; a serial loop adds every tool's time instead. The controls themselves cost nothing per call: loop detection is a dictionary of (tool, arguments) counts, and a budget is a comparison.

What breaks

  • The model never says "done". Without a step limit the loop runs until the money does. Set max_steps, and treat hitting it as an outcome to monitor, not an error to hide.
  • The same call, again and again. A model retrying one search burns a call per repeat and never progresses. Count (tool, canonical arguments) and stop at a threshold.
  • A half-written tool call. A reply cut off by max_tokens can end mid-argument. Check the stop reason before running anything; a truncated call must never execute.
  • A tool throws. Crashing the loop discards the work so far; swallowing the error makes the model guess. Return the failure as an error result with a message the model can act on ("as_of must be YYYY-MM-DD"); in the lesson, the model fixes its own call on the next step.
  • A tool the model made up. An unknown name is a KeyError in a naive loop. Answer with an error result listing the real tools.
  • Results split across messages. The API pairs each request and result by id inside one user turn; splitting them breaks the pairing and teaches the model to stop making parallel calls.
  • Guessing outside its authority. A refund over the limit needs a person. Make hand-off a tool, so it ends the run with a logged reason instead of a confident wrong answer.

In the wild

ReAct (Yao et al., 2022) named the pattern the loop implements: a short piece of reasoning, an action, an observation, repeated. Claude's SDKs ship a tool runner that loops until the model returns no tool use or max_iterations is reached, and their docs send you to the manual loop when you need approval or custom logging. The OpenAI Agents SDK runs the same call, classify, run-tools cycle and raises MaxTurnsExceeded past max_turns. LangGraph bounds a graph with a recursion limit and raises GraphRecursionError when it trips. Anthropic's Building effective agents adds the operational advice: agents trade cost and latency for open-ended capability, so test them in sandboxes and put guardrails in. Every one of these products is this lesson's diamond ("done, or budget, or step limit?") with a different name on it.

Go deeper

Level 2 traces one round trip message by message, orders the seven exits the way the loop checks them and says why that order matters, fans tool calls out and back in, and derives the quadratic cost with a formula you can rerun against the budget-burner demo. If you only needed to choose, you are done.

Level 2

How it works, from scratch

The loop itself is ten lines: call the model, run the tools it asks for, append the results, repeat until it stops asking. Everything else in this lesson exists because the ten-line version fails in production, and each failure gets a named control.

Chapter 1

The idea

Everyday picture A new assistant runs errands for you. They can't do anything themselves. They come back after every errand and say: "here's what I found; next I'd like to do X." You carry out X and hand them the result. They decide the next step, and so on, until they say "done" (or you say "that's enough, you've spent the budget"). The assistant is the model, the errands are tool calls, and you are the loop in this file.

Tiny worked example "How many PTO days does alice have left, and what rolls over?" Here's the real transcript the happy-path demo produces:

[user]       How many PTO days does alice have left, and what rolls over?
[assistant]  I'll check the PTO policy and alice's balance at the same time.
             tool_use toolu_1 search_kb       {"query": "PTO rollover"}
             tool_use toolu_2 get_pto_balance {"employee": "alice", "as_of": "2026-09-25"}
[user]       tool_result toolu_1  "[hr-001] Paid time off: ... Unused PTO up to 5 days rolls over. ..."
             tool_result toolu_2  {"employee": "alice", "as_of": "2026-09-25", "days_left": 12}
[assistant]  Alice has 12 PTO days left. Up to 5 unused days roll over to next year [hr-001].
             (stop_reason: end_turn)

Two model calls, two tools run in parallel, one answer. ReAct ("reason + act") is the research name for this pattern of alternating a bit of reasoning ("I'll check both at once") with actions (tool calls). The loop itself is ten lines of code. Everything else in this file exists because the ten-line version fails in production.

Figure 1 · Diagram

Reading it: start at Goal on the left and follow the arrows around the cycle. The model only ever does the Think box: it looks at everything so far and picks the next action. Call a tool and Observe result are your code. The diamond is the only exit, and it has three ways out: the model says it's done, a limit trips, or the agent hands the task to a person. An agent without that diamond is a loop you can't stop.

In code: run_agent is this loop, and it returns an AgentResult holding the final text, the stop cause, every step, the usage and the cost. search_kb and get_pto_balance are the two tools in the example.

Chapter 2

One round trip, message by message

Figure 2 · Diagram

Reading it: time runs top to bottom. The model's reply is a request (tool_use), never an action. Everything between that request and the next call to the model happens in your loop, which is where every control in this file lives: validation, budgets, loop detection, concurrency. Notice the model is called twice for one question. Each call re-sends the whole history, and that's where the cost comes from.

In code: each pass through run_agent appends the model's full assistant content, runs the requested tools, and appends every result in one user message; a Step records that one model call and the ToolExecutions it triggered.

Chapter 3

The controls, in the order the loop checks them

Everyday picture Lending your car to a new driver: a full tank but no more fuel money (budget), "come back by six" (step limit), "if you drive past the same petrol station three times, you're lost, so come home" (loop detection), and "call me if anything feels off" (hand-off to a human).

Figure 3 · Diagram

Reading it: every box that starts with stop: is an exit, and each of the seven is a named stop_cause in AgentResult. Read top to bottom: first trust the model's own stop reason, then protect your wallet (budget), then honor an explicit escalation, then protect against repetition, and only then spend time running tools. The order matters. A truncated response (max_tokens) is checked before tools run, because a call cut off mid-argument must never execute.
Failure What happens Control in run_agent
Model never says "done" Runs forever max_steps
Each turn re-sends the whole conversation Cost grows quadratically with steps max_total_tokens, max_cost_usd
Model repeats the same call Burns money, no progress loop detection on (tool, canonical args)
Tool raises Loop crashes, or the model never learns why errors become is_error tool results with actionable text
Unknown tool / hallucinated name KeyError is_error result listing the real tools
Several independent calls in one turn Slow if run one by one run concurrently, return all results in one user message
Output cut off (max_tokens) Half-written tool call stop with cause max_tokens
Safety refusal (refusal) Content is not an answer stop with cause refusal
Task is outside the agent's authority Agent guesses a handoff_to_human tool that ends the run cleanly

"It stopped" is not an outcome you can monitor; "it stopped because of loop_detected at step 3" is.

In code: AgentConfig holds every limit (steps, tokens, dollars, the loop threshold) and the prices AgentConfig.cost uses to turn primer.agents.llm.Usage into dollars. A tool raises ToolError to send the model an actionable error result, and HANDOFF_TOOL_DEF is the hand-off tool's definition.

Chapter 4

Parallel tool calls: fan out, fan in

Everyday picture Three errands in three different shops. You can do them one after another, or send three friends at once and be done when the slowest one gets back.

Figure 4 · Diagram

Reading it: the model asked for A, B and C in the same turn, before seeing any result, so they can't depend on each other and it's safe to run them at once. Wall-clock time becomes the slowest call instead of the sum. On the way back, all three results go in a single user message. The API pairs each result with its request by id, and splitting them teaches the model to stop asking for parallel calls.

In code: execute_tools runs one turn's calls on a thread pool and returns their results in the order they were asked for, turning an unknown tool name or a raised exception into an error result instead of a crash.

Chapter 5

Why cost grows quadratically

Each call sends the entire conversation so far. If every step adds about tokens, call sends roughly input tokens, so a run of steps sends

Level 3: the formula and its symbols

Symbols

Symbol Meaning
new tokens each step adds to the conversation (tool call + result)
the step number, 1, 2, 3, ...
total steps in the run

In words: step re-sends everything from the steps before it, so the total is times , which is about half of squared times .

On the example: with and , the total is input tokens, not the you'd guess. Double to and it's , almost 4x.

In Python:

t = 500
def total_input(n):
    # Σ_k k·t: step k re-sends k steps' worth
    return sum(k * t for k in range(1, n + 1))
total_input(10)  # → 27500
total_input(20)  # → 105000
# almost 4x
round(total_input(20) / total_input(10), 1)  # → 3.8

Figure 5 · Drawn from the lesson's code

0 2 4 6 8 10 step 0 2000 4000 6000 8000 10000 12000 tokens A runaway agent: each call re-sends the whole history running total budget (12,000) tokens sent this step

Each call of a runaway agent sends more tokens than the last, and the run stops at the first call where the running total passes the 12,000-token budget

Reading it: each bar is one call to the model in the budget-burner demo (a model that keeps searching). Bars grow step by step because the history grows. The line is the running total, which is what you pay for, and it bends upward. The dashed horizontal line is the token budget. The run stops at the first step where the total crosses it, instead of carrying on.

Figure 6 · Drawn from the lesson's code

0 5 10 15 20 25 30 steps in the run 0 50000 100000 150000 200000 total input tokens Cumulative input tokens (500 new tokens per step) if each step cost the same (linear) re-sending history (quadratic)

At 10 steps re-sending the history costs 27,500 input tokens against the 5,000 a flat per-step cost suggests, and the gap keeps widening

Reading it: the x-axis is the number of steps in a run, and the y-axis is the total input tokens sent. The straight line is what people intuitively expect ("each step costs the same"). The curve is what actually happens when every call re-sends the history. At 10 steps the gap is about 5x, and it keeps widening. That gap is why step limits, trimming tool output, and prompt caching (primer.agents.cost) matter.

Figure 7 · Drawn from the lesson's code

1 2 3 4 5 6 7 8 tool calls requested in one turn 1 2 3 4 5 wall-clock seconds Parallel tool calls: time is the slowest call, not the sum sequential (sum of latencies) concurrent (max latency)

Run one after another, tool calls take the sum of their latencies and keep climbing; run concurrently, they take only as long as the slowest call

Reading it: the x-axis is how many tools the model requested in one turn, each with a realistic latency between 0.2 s and 1.2 s. Run one after another, the time is the sum and climbs with every tool. Run concurrently, it's the max, which flattens out near the slowest single call. Parallel tool calls are among the cheapest latency wins in agent systems.

In code: cumulative_input_tokens evaluates the formula above, times , for any run length.

Chapter 6

Running it against a real model

The loop only depends on the LLM protocol, so the same code runs against Claude:

from primer.agents.llm import ClaudeLLM
from primer.agents.agent_loop import run_agent, AgentConfig

result = run_agent(
    ClaudeLLM(),                      # needs `pip install anthropic` + credentials
    task="How many PTO days does alice have left, and what's the rollover rule?",
    tools=HR_TOOLS,                   # name -> python function
    tool_defs=HR_TOOL_DEFS,           # Anthropic tool definitions (JSON Schema)
    system="You are an HR assistant. Use tools; cite doc ids.",
    config=AgentConfig(max_steps=8, max_cost_usd=0.50),
)
print(result.stop_cause, result.final_text)

The SDK also ships a tool runner (client.beta.messages.tool_runner) that drives this loop for you. Owning the loop, as here, is what you do when you need custom budgets, loop detection, approval gates or tracing in exactly your shape.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: An agent in production suddenly costs 5x more per task. What do you check first?Think it through, then reveal

A: Steps per task and tokens per step, from traces. A jump usually means a loop (the same call repeated), a tool that started returning huge payloads, or a prompt change that stopped the model from recognizing "done". Step budgets and loop detection cap the damage; alerts on tokens-per-task catch it early.

Question 2Q: A tool throws an exception mid-run. What should the loop do?Think it through, then reveal

A: Catch it and return a tool_result with is_error: true and a message the model can act on ("date must be YYYY-MM-DD, got 'next friday'"). The model usually fixes its next call. Crashing the loop throws away the work so far; swallowing the error silently makes the model guess.

Question 3Q: Why return all parallel tool results in one message?Think it through, then reveal

A: The API pairs each tool_use with its tool_result by id in the next user turn. Splitting results across several messages breaks that pairing and teaches the model to stop making parallel calls, which slows every run.

Question 4Q: When should an agent hand off to a human?Think it through, then reveal

A: When it lacks authority (refunds over a limit), lacks information after reasonable search, detects conflicting sources, or hits a budget. Make hand-off a tool, so it is an explicit, logged outcome rather than a vague final answer.

Primary sources

The papers behind this lesson

Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022).

It showed that interleaving short reasoning traces with tool actions, then observing the results, beats reasoning alone or acting alone, and this loop is the shape of nearly every agent built since.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
  • Claude tool use, implementing the loop: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
  • Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022): https://arxiv.org/abs/2210.03629
  • Lilian Weng, LLM Powered Autonomous Agents: https://lilianweng.github.io/posts/2023-06-23-agent/

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.