rumblr Work in progressWIP

● The AI Primer · Lesson 40 · Part 2: building systems people rely on

The agent loop, production-grade

This lesson covers A production agent loop: budgets, loop detection, recovery

Members · open during launch 19 min7 figures and diagrams
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. The agent loop is: call model, run the tools it asks for, append results, repeat until end_turn.
  2. The model never executes anything. Your code validates and runs every tool call.
  3. Production loops need step limits, token/cost budgets, loop detection, error-as-result handling, and a human hand-off.
  4. Parallel tool calls: run them concurrently, return all results in a single user message.
  5. Input cost grows roughly with the square of the number of steps, because every call re-sends the history.

Level 2

How it works, from scratch

The loop itself is ten lines: call the model, run the tools it asks for, append the results, repeat until it stops asking. Everything else in this lesson exists because the ten-line version fails in production, and each failure gets a named control.

Chapter 1

The idea

Everyday picture A new assistant runs errands for you. They can't do anything themselves. They come back after every errand and say: "here's what I found; next I'd like to do X." You carry out X and hand them the result. They decide the next step, and so on, until they say "done" (or you say "that's enough, you've spent the budget"). The assistant is the model, the errands are tool calls, and you are the loop in this file.

Tiny worked example "How many PTO days does alice have left, and what rolls over?" Here's the real transcript the happy-path demo produces:

[user]       How many PTO days does alice have left, and what rolls over?
[assistant]  I'll check the PTO policy and alice's balance at the same time.
             tool_use toolu_1 search_kb       {"query": "PTO rollover"}
             tool_use toolu_2 get_pto_balance {"employee": "alice", "as_of": "2026-09-25"}
[user]       tool_result toolu_1  "[hr-001] Paid time off: ... Unused PTO up to 5 days rolls over. ..."
             tool_result toolu_2  {"employee": "alice", "as_of": "2026-09-25", "days_left": 12}
[assistant]  Alice has 12 PTO days left. Up to 5 unused days roll over to next year [hr-001].
             (stop_reason: end_turn)

Two model calls, two tools run in parallel, one answer. ReAct ("reason + act") is the research name for this pattern of alternating a bit of reasoning ("I'll check both at once") with actions (tool calls). The loop itself is ten lines of code. Everything else in this file exists because the ten-line version fails in production.

Figure 1 · Diagram

Reading it: start at Goal on the left and follow the arrows around the cycle. The model only ever does the Think box: it looks at everything so far and picks the next action. Call a tool and Observe result are your code. The diamond is the only exit, and it has three ways out: the model says it's done, a limit trips, or the agent hands the task to a person. An agent without that diamond is a loop you can't stop.

In code: run_agent is this loop, and it returns an AgentResult holding the final text, the stop cause, every step, the usage and the cost. search_kb and get_pto_balance are the two tools in the example.

Chapter 2

One round trip, message by message

Figure 2 · Diagram

Reading it: time runs top to bottom. The model's reply is a request (tool_use), never an action. Everything between that request and the next call to the model happens in your loop, which is where every control in this file lives: validation, budgets, loop detection, concurrency. Notice the model is called twice for one question. Each call re-sends the whole history, and that's where the cost comes from.

In code: each pass through run_agent appends the model's full assistant content, runs the requested tools, and appends every result in one user message; a Step records that one model call and the ToolExecutions it triggered.

Chapter 3

The controls, in the order the loop checks them

Everyday picture Lending your car to a new driver: a full tank but no more fuel money (budget), "come back by six" (step limit), "if you drive past the same petrol station three times, you're lost, so come home" (loop detection), and "call me if anything feels off" (hand-off to a human).

Figure 3 · Diagram

Reading it: every box that starts with stop: is an exit, and each of the seven is a named stop_cause in AgentResult. Read top to bottom: first trust the model's own stop reason, then protect your wallet (budget), then honor an explicit escalation, then protect against repetition, and only then spend time running tools. The order matters. A truncated response (max_tokens) is checked before tools run, because a call cut off mid-argument must never execute.
Failure What happens Control in run_agent
Model never says "done" Runs forever max_steps
Each turn re-sends the whole conversation Cost grows quadratically with steps max_total_tokens, max_cost_usd
Model repeats the same call Burns money, no progress loop detection on (tool, canonical args)
Tool raises Loop crashes, or the model never learns why errors become is_error tool results with actionable text
Unknown tool / hallucinated name KeyError is_error result listing the real tools
Several independent calls in one turn Slow if run one by one run concurrently, return all results in one user message
Output cut off (max_tokens) Half-written tool call stop with cause max_tokens
Safety refusal (refusal) Content is not an answer stop with cause refusal
Task is outside the agent's authority Agent guesses a handoff_to_human tool that ends the run cleanly

"It stopped" is not an outcome you can monitor; "it stopped because of loop_detected at step 3" is.

In code: AgentConfig holds every limit (steps, tokens, dollars, the loop threshold) and the prices AgentConfig.cost uses to turn primer.agents.llm.Usage into dollars. A tool raises ToolError to send the model an actionable error result, and HANDOFF_TOOL_DEF is the hand-off tool's definition.

Chapter 4

Parallel tool calls: fan out, fan in

Everyday picture Three errands in three different shops. You can do them one after another, or send three friends at once and be done when the slowest one gets back.

Figure 4 · Diagram

Reading it: the model asked for A, B and C in the same turn, before seeing any result, so they can't depend on each other and it's safe to run them at once. Wall-clock time becomes the slowest call instead of the sum. On the way back, all three results go in a single user message. The API pairs each result with its request by id, and splitting them teaches the model to stop asking for parallel calls.

In code: execute_tools runs one turn's calls on a thread pool and returns their results in the order they were asked for, turning an unknown tool name or a raised exception into an error result instead of a crash.

Chapter 5

Why cost grows quadratically

Each call sends the entire conversation so far. If every step adds about tokens, call sends roughly input tokens, so a run of steps sends

Level 3: the formula and its symbols

Symbols

Symbol Meaning
new tokens each step adds to the conversation (tool call + result)
the step number, 1, 2, 3, ...
total steps in the run

In words: step re-sends everything from the steps before it, so the total is times , which is about half of squared times .

On the example: with and , the total is input tokens, not the you'd guess. Double to and it's , almost 4x.

In Python:

t = 500
def total_input(n):
    # Σ_k k·t: step k re-sends k steps' worth
    return sum(k * t for k in range(1, n + 1))
total_input(10)  # → 27500
total_input(20)  # → 105000
# almost 4x
round(total_input(20) / total_input(10), 1)  # → 3.8

Figure 5 · Chart

0 2 4 6 8 10 step 0 2000 4000 6000 8000 10000 12000 tokens A runaway agent: each call re-sends the whole history running total budget (12,000) tokens sent this step

Each call of a runaway agent sends more tokens than the last, and the run stops at the first call where the running total passes the 12,000-token budget

Reading it: each bar is one call to the model in the budget-burner demo (a model that keeps searching). Bars grow step by step because the history grows. The line is the running total, which is what you pay for, and it bends upward. The dashed horizontal line is the token budget. The run stops at the first step where the total crosses it, instead of carrying on.

Figure 6 · Chart

0 5 10 15 20 25 30 steps in the run 0 50000 100000 150000 200000 total input tokens Cumulative input tokens (500 new tokens per step) if each step cost the same (linear) re-sending history (quadratic)

At 10 steps re-sending the history costs 27,500 input tokens against the 5,000 a flat per-step cost suggests, and the gap keeps widening

Reading it: the x-axis is the number of steps in a run, and the y-axis is the total input tokens sent. The straight line is what people intuitively expect ("each step costs the same"). The curve is what actually happens when every call re-sends the history. At 10 steps the gap is about 5x, and it keeps widening. That gap is why step limits, trimming tool output, and prompt caching (primer.agents.cost) matter.

Figure 7 · Chart

1 2 3 4 5 6 7 8 tool calls requested in one turn 1 2 3 4 5 wall-clock seconds Parallel tool calls: time is the slowest call, not the sum sequential (sum of latencies) concurrent (max latency)

Run one after another, tool calls take the sum of their latencies and keep climbing; run concurrently, they take only as long as the slowest call

Reading it: the x-axis is how many tools the model requested in one turn, each with a realistic latency between 0.2 s and 1.2 s. Run one after another, the time is the sum and climbs with every tool. Run concurrently, it's the max, which flattens out near the slowest single call. Parallel tool calls are among the cheapest latency wins in agent systems.

In code: cumulative_input_tokens evaluates the formula above, times , for any run length.

Chapter 6

Running it against a real model

The loop only depends on the LLM protocol, so the same code runs against Claude:

from primer.agents.llm import ClaudeLLM
from primer.agents.agent_loop import run_agent, AgentConfig

result = run_agent(
    ClaudeLLM(),                      # needs `pip install anthropic` + credentials
    task="How many PTO days does alice have left, and what's the rollover rule?",
    tools=HR_TOOLS,                   # name -> python function
    tool_defs=HR_TOOL_DEFS,           # Anthropic tool definitions (JSON Schema)
    system="You are an HR assistant. Use tools; cite doc ids.",
    config=AgentConfig(max_steps=8, max_cost_usd=0.50),
)
print(result.stop_cause, result.final_text)

The SDK also ships a tool runner (client.beta.messages.tool_runner) that drives this loop for you. Owning the loop, as here, is what you do when you need custom budgets, loop detection, approval gates or tracing in exactly your shape.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: An agent in production suddenly costs 5x more per task. What do you check first?Think it through, then reveal

A: Steps per task and tokens per step, from traces. A jump usually means a loop (the same call repeated), a tool that started returning huge payloads, or a prompt change that stopped the model from recognizing "done". Step budgets and loop detection cap the damage; alerts on tokens-per-task catch it early.

Question 2Q: A tool throws an exception mid-run. What should the loop do?Think it through, then reveal

A: Catch it and return a tool_result with is_error: true and a message the model can act on ("date must be YYYY-MM-DD, got 'next friday'"). The model usually fixes its next call. Crashing the loop throws away the work so far; swallowing the error silently makes the model guess.

Question 3Q: Why return all parallel tool results in one message?Think it through, then reveal

A: The API pairs each tool_use with its tool_result by id in the next user turn. Splitting results across several messages breaks that pairing and teaches the model to stop making parallel calls, which slows every run.

Question 4Q: When should an agent hand off to a human?Think it through, then reveal

A: When it lacks authority (refunds over a limit), lacks information after reasonable search, detects conflicting sources, or hits a budget. Make hand-off a tool, so it is an explicit, logged outcome rather than a vague final answer.

Primary sources

The papers behind this lesson

Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022).

It showed that interleaving short reasoning traces with tool actions, then observing the results, beats reasoning alone or acting alone, and this loop is the shape of nearly every agent built since.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
  • Claude tool use, implementing the loop: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
  • Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022): https://arxiv.org/abs/2210.03629
  • Lilian Weng, LLM Powered Autonomous Agents: https://lilianweng.github.io/posts/2023-06-23-agent/

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.