At a glance
Key takeaways
- The agent loop is: call model, run the tools it asks for, append results, repeat until
end_turn. - The model never executes anything. Your code validates and runs every tool call.
- Production loops need step limits, token/cost budgets, loop detection, error-as-result handling, and a human hand-off.
- Parallel tool calls: run them concurrently, return all results in a single user message.
- Input cost grows roughly with the square of the number of steps, because every call re-sends the history.
Level 2
How it works, from scratch
The loop itself is ten lines: call the model, run the tools it asks for, append the results, repeat until it stops asking. Everything else in this lesson exists because the ten-line version fails in production, and each failure gets a named control.
Chapter 1
The idea
Everyday picture A new assistant runs errands for you. They can't do anything themselves. They come back after every errand and say: "here's what I found; next I'd like to do X." You carry out X and hand them the result. They decide the next step, and so on, until they say "done" (or you say "that's enough, you've spent the budget"). The assistant is the model, the errands are tool calls, and you are the loop in this file.
Tiny worked example "How many PTO days does alice have left, and what rolls over?" Here's the real transcript the happy-path demo produces:
[user] How many PTO days does alice have left, and what rolls over?
[assistant] I'll check the PTO policy and alice's balance at the same time.
tool_use toolu_1 search_kb {"query": "PTO rollover"}
tool_use toolu_2 get_pto_balance {"employee": "alice", "as_of": "2026-09-25"}
[user] tool_result toolu_1 "[hr-001] Paid time off: ... Unused PTO up to 5 days rolls over. ..."
tool_result toolu_2 {"employee": "alice", "as_of": "2026-09-25", "days_left": 12}
[assistant] Alice has 12 PTO days left. Up to 5 unused days roll over to next year [hr-001].
(stop_reason: end_turn)
Two model calls, two tools run in parallel, one answer. ReAct ("reason + act") is the research name for this pattern of alternating a bit of reasoning ("I'll check both at once") with actions (tool calls). The loop itself is ten lines of code. Everything else in this file exists because the ten-line version fails in production.
Figure 1 · Diagram
flowchart LR
G[Goal] --> T[Think<br/>choose next step]
T --> C[Call a tool]
C --> O[Observe result]
O --> D{Done, or budget<br/>or step limit hit?}
D -->|No| T
D -->|Yes| F[Final answer<br/>or hand to human]
In code: run_agent is this loop, and it returns an AgentResult
holding the final text, the stop cause, every step, the usage and the cost.
search_kb and get_pto_balance are the two tools in the example.
Chapter 2
One round trip, message by message
Figure 2 · Diagram
sequenceDiagram participant U as User participant L as Your loop participant M as Model participant T as Tools U->>L: task L->>M: system + messages + tool definitions M-->>L: tool_use blocks (stop_reason = tool_use) L->>L: validate, check budget, check for loops L->>T: run the calls (concurrently) T-->>L: results or errors L->>M: ONE user message with every tool_result M-->>L: final text (stop_reason = end_turn) L-->>U: answer + stop cause
tool_use), never an action. Everything between that request and the next
call to the model happens in your loop, which is where every control in
this file lives: validation, budgets, loop detection, concurrency. Notice
the model is called twice for one question. Each call re-sends the whole
history, and that's where the cost comes from.In code: each pass through run_agent appends the model's full
assistant content, runs the requested tools, and appends every result in one
user message; a Step records that one model call and the ToolExecutions
it triggered.
Chapter 3
The controls, in the order the loop checks them
Everyday picture Lending your car to a new driver: a full tank but no more fuel money (budget), "come back by six" (step limit), "if you drive past the same petrol station three times, you're lost, so come home" (loop detection), and "call me if anything feels off" (hand-off to a human).
Figure 3 · Diagram
flowchart TD
R[Model response] --> S1{stop_reason?}
S1 -->|refusal| X1[stop: refusal]
S1 -->|max_tokens| X2[stop: max_tokens<br/>don't run half-written calls]
S1 -->|end_turn| X3[stop: completed]
S1 -->|tool_use| B{Over token or<br/>dollar budget?}
B -->|yes| X4[stop: budget_exceeded]
B -->|no| H{Called handoff_to_human?}
H -->|yes| X5[stop: handoff]
H -->|no| L{Same tool + same args<br/>seen N times?}
L -->|yes| X6[stop: loop_detected]
L -->|no| E[Run tools concurrently]
E --> A[Append all results<br/>in one user message]
A --> N{Step limit hit?}
N -->|yes| X7[stop: max_steps]
N -->|no| M[Call the model again]
stop: is an exit, and each of
the seven is a named stop_cause in AgentResult.
Read top to bottom: first trust the model's own stop reason, then protect
your wallet (budget), then honor an explicit escalation, then protect
against repetition, and only then spend time running tools. The order
matters. A truncated response (max_tokens) is checked before tools run,
because a call cut off mid-argument must never execute.| Failure | What happens | Control in run_agent |
|---|---|---|
| Model never says "done" | Runs forever | max_steps |
| Each turn re-sends the whole conversation | Cost grows quadratically with steps | max_total_tokens, max_cost_usd |
| Model repeats the same call | Burns money, no progress | loop detection on (tool, canonical args) |
| Tool raises | Loop crashes, or the model never learns why | errors become is_error tool results with actionable text |
| Unknown tool / hallucinated name | KeyError |
is_error result listing the real tools |
| Several independent calls in one turn | Slow if run one by one | run concurrently, return all results in one user message |
Output cut off (max_tokens) |
Half-written tool call | stop with cause max_tokens |
Safety refusal (refusal) |
Content is not an answer | stop with cause refusal |
| Task is outside the agent's authority | Agent guesses | a handoff_to_human tool that ends the run cleanly |
"It stopped" is not an outcome you can monitor; "it stopped because of loop_detected at step 3" is.
In code: AgentConfig holds every limit (steps, tokens, dollars, the
loop threshold) and the prices AgentConfig.cost uses to turn primer.agents.llm.Usage into
dollars. A tool raises ToolError to send the model an actionable error
result, and HANDOFF_TOOL_DEF is the hand-off tool's definition.
Chapter 4
Parallel tool calls: fan out, fan in
Everyday picture Three errands in three different shops. You can do them one after another, or send three friends at once and be done when the slowest one gets back.
Figure 4 · Diagram
flowchart LR M[Assistant turn<br/>tool_use A, tool_use B, tool_use C] --> A[run A] M --> B[run B] M --> C[run C] A --> J[One user message:<br/>tool_result A, B, C<br/>matched by tool_use_id] B --> J C --> J
In code: execute_tools runs one turn's calls on a thread pool and
returns their results in the order they were asked for, turning an unknown
tool name or a raised exception into an error result instead of a crash.
Chapter 5
Why cost grows quadratically
Each call sends the entire conversation so far. If every step adds about tokens, call sends roughly input tokens, so a run of steps sends
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning |
|---|---|
| new tokens each step adds to the conversation (tool call + result) | |
| the step number, 1, 2, 3, ... | |
| total steps in the run |
In words: step re-sends everything from the steps before it, so the total is times , which is about half of squared times .
On the example: with and , the total is input tokens, not the you'd guess. Double to and it's , almost 4x.
In Python:
t = 500
def total_input(n):
# Σ_k k·t: step k re-sends k steps' worth
return sum(k * t for k in range(1, n + 1))
total_input(10) # → 27500
total_input(20) # → 105000
# almost 4x
round(total_input(20) / total_input(10), 1) # → 3.8
Figure 5 · Chart
Each call of a runaway agent sends more tokens than the last, and the run stops at the first call where the running total passes the 12,000-token budget
Figure 6 · Chart
At 10 steps re-sending the history costs 27,500 input tokens against the 5,000 a flat per-step cost suggests, and the gap keeps widening
primer.agents.cost) matter.Figure 7 · Chart
Run one after another, tool calls take the sum of their latencies and keep climbing; run concurrently, they take only as long as the slowest call
In code: cumulative_input_tokens evaluates the formula above,
times , for any run length.
Chapter 6
Running it against a real model
The loop only depends on the LLM protocol, so the same code runs against
Claude:
from primer.agents.llm import ClaudeLLM
from primer.agents.agent_loop import run_agent, AgentConfig
result = run_agent(
ClaudeLLM(), # needs `pip install anthropic` + credentials
task="How many PTO days does alice have left, and what's the rollover rule?",
tools=HR_TOOLS, # name -> python function
tool_defs=HR_TOOL_DEFS, # Anthropic tool definitions (JSON Schema)
system="You are an HR assistant. Use tools; cite doc ids.",
config=AgentConfig(max_steps=8, max_cost_usd=0.50),
)
print(result.stop_cause, result.final_text)
The SDK also ships a tool runner (client.beta.messages.tool_runner)
that drives this loop for you. Owning the loop, as here, is what you do
when you need custom budgets, loop detection, approval gates or tracing in
exactly your shape.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Q: An agent in production suddenly costs 5x more per task. What do you check first?Think it through, then reveal
A: Steps per task and tokens per step, from traces. A jump usually means a loop (the same call repeated), a tool that started returning huge payloads, or a prompt change that stopped the model from recognizing "done". Step budgets and loop detection cap the damage; alerts on tokens-per-task catch it early.
Question 2Q: A tool throws an exception mid-run. What should the loop do?Think it through, then reveal
A: Catch it and return a tool_result with is_error: true and a message the
model can act on ("date must be YYYY-MM-DD, got 'next friday'"). The model
usually fixes its next call. Crashing the loop throws away the work so far;
swallowing the error silently makes the model guess.
Question 3Q: Why return all parallel tool results in one message?Think it through, then reveal
A: The API pairs each tool_use with its tool_result by id in the next user
turn. Splitting results across several messages breaks that pairing and
teaches the model to stop making parallel calls, which slows every run.
Question 4Q: When should an agent hand off to a human?Think it through, then reveal
A: When it lacks authority (refunds over a limit), lacks information after reasonable search, detects conflicting sources, or hits a budget. Make hand-off a tool, so it is an explicit, logged outcome rather than a vague final answer.
Primary sources
The papers behind this lesson
It showed that interleaving short reasoning traces with tool actions, then observing the results, beats reasoning alone or acting alone, and this loop is the shape of nearly every agent built since.
Read the annotated companion →The paper ↗Researcher's shelf
Further reading
- Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
- Claude tool use, implementing the loop: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022): https://arxiv.org/abs/2210.03629
- Lilian Weng, LLM Powered Autonomous Agents: https://lilianweng.github.io/posts/2023-06-23-agent/
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.