rumblr Work in progressWIP

● The AI Primer · Lesson 38 · Part 2: building systems people rely on

Talking to a model

messages, and what tool calling really is

This lesson covers The message format, and what tool calling really is

Members · open during launch 16 min3 figures and diagrams
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. A model call sends the full list of messages and returns one message; the model keeps no state between calls.
  2. A tool call is a structured request (tool_use). Your code validates it, runs it, and returns a tool_result with the matching id. The model executes nothing.
  3. stop_reason tells you what to do next: tool_use means run tools and call again, end_turn means done.
  4. Every call re-sends the whole history, so input cost grows quadratically with the number of steps in a loop.

Level 2

How it works, from scratch

Every applied-AI system in this primer, from a one-shot classifier to a multi-agent research team, talks to the model the same way: it sends a list of messages and gets one message back. This lesson shows that exchange exactly, and then the one feature that turns a chatbot into an agent: tool calling.

Chapter 1

The everyday picture

Think of the model as a brilliant consultant who works by post. You mail a letter with your whole conversation so far (the consultant keeps no notes between letters) and get one letter back. The reply is either an answer, or a filled-in request form: "please look up the weather in Paris and send me the result." The consultant never picks up the phone themselves. You decide whether to run the request, run it, and mail back the result, together with the whole conversation again.

Two facts fall straight out of this picture, and they explain most of the behaviour of real systems:

  1. The model executes nothing. A tool call is a request. Your code is the only thing that acts, so your code is where safety lives.
  2. Every call resends everything. The model has no memory between calls, so each request carries the full history. Long conversations cost more on every single turn.

Chapter 2

A tiny worked example: one tool call, traced

A user asks "What's the weather in Paris?" and the application offers one tool, get_weather. Four messages later, the user has an answer:

# Role Content Who produced it
1 user "What's the weather in Paris?" the user
2 assistant tool_use id=toolu_1, name=get_weather, input={"city": "Paris"} the model (1st call)
3 user tool_result for toolu_1: "18°C, sunny" your code, after running the tool
4 assistant "It's 18°C and sunny in Paris." the model (2nd call)

Message 3 has role user even though no human typed it: tool results always travel back in the user's turn. The id in message 2 and the tool_use_id in message 3 match, which is how the model knows which request a result answers when it asked for several at once. trace_tool_call_round_trip() produces exactly this transcript.

Figure 1 · Diagram

Reading it: time runs downwards. Solid arrows are requests and dashed arrows are replies. Notice that the model never talks to the tool: every arrow into get_weather starts at your code, and the self-arrow on your code ("validate, check permissions") is where production systems put their guardrails. Notice too that the second call to the model carries messages 1 to 3, not just the new result: the model is stateless, so the history travels every time.

In code: ToolCall holds the request in message 2, and tool_result_block builds the reply in message 3, carrying the matching id.

Chapter 3

The message format

Messages use the Anthropic Messages API shape directly, so what you learn here maps one-to-one onto production code:

{"role": "user", "content": "What's our PTO policy?"}
{"role": "assistant", "content": [
    {"type": "text", "text": "Let me look that up."},
    {"type": "tool_use", "id": "toolu_1", "name": "search_kb", "input": {"query": "PTO"}},
]}
{"role": "user", "content": [
    {"type": "tool_result", "tool_use_id": "toolu_1", "content": "hr-001: ..."},
]}
Field Meaning
role user (the human, or your code returning tool results) or assistant (the model)
content a plain string, or a list of content blocks
text block ordinary words
tool_use block the model asking for a tool: id, name, and input (a JSON object matching the tool's schema)
tool_result block your answer to one tool_use, matched by tool_use_id; set is_error: true when the tool failed
stop_reason why the reply ended: end_turn (finished), tool_use (wants a tool), max_tokens (ran out of room), refusal (declined)

A tool definition is a name, a description and a JSON Schema for the input. The model chooses tools by reading those descriptions, so writing them well is prompt engineering (see primer.agents.tools).

In code: LLMResponse is one reply, normalized: its text, its ToolCall list, its stop reason, its Usage and the exact content blocks to append as the assistant turn. content_blocks, last_user_text, tool_results and tool_calls_so_far read a conversation in this format.

Chapter 4

Two implementations of one interface

Every agent lesson is written against one small interface, LLM, with one method, LLM.complete(system=..., messages=..., tools=...). Two classes implement it:

Figure 2 · Diagram

Reading it: the agent code on the right-hand side never knows which box it's talking to. ScriptedLLM lets every lesson run offline and deterministically, and lets tests stage a model that loops, calls the wrong tool or follows an injected instruction, which is hard to get a real model to do on cue. Swapping in ClaudeLLM runs the identical loop against the real thing.

In code: ClaudeLLM.complete builds its request with claude_request and turns the API's reply into an LLMResponse. OllamaLLM is a third implementation for a local open model (text only), using ollama_request and parse_ollama_reply.

Chapter 5

Cost: why every call pays for the whole conversation

Because the model is stateless, input tokens are billed per call on the entire history. In an agent loop with calls, where each call adds about new tokens to a history that started at tokens, the total input billed is:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
number of model calls in the loop 10
tokens in the first request (system prompt, tools, question) 2,000
tokens each step adds (the model's tool call plus the tool's result) 500
which call we're on, 1 to
add up the cost of every call
0 + 1 + … + (n − 1): how many "steps of history" pile up in total 45

In words: "each call pays for the starting prompt plus everything added so far, so the history term grows with the square of the number of steps."

With the numbers: 10 × 2,000 + 500 × 45 = 20,000 + 22,500 = 42,500 input tokens for a task whose final conversation is only 7,000 tokens long: the tenth call re-sends 6,500 tokens, and its own 500-token step brings the history to 7,000.

Level 3: in Python
n, h_0, t = 10, 2000, 500
# call i re-sends h_0 and i - 1 steps
calls = [h_0 + (i - 1) * t for i in range(1, n + 1)]
calls[0], calls[-1]  # → (2000, 6500)
# Σ over every call
sum(calls)  # → 42500
# the shortcut on the right agrees
n * h_0 + t * n * (n - 1) // 2  # → 42500

Figure 3 · Drawn from the lesson's code

2 4 6 8 10 model call in the loop 0 5000 10000 15000 20000 25000 30000 35000 40000 input tokens Every call re-sends the whole history running total final conversation size (billed once?) input tokens billed on this call

Ten calls bill 42,500 input tokens in total, six times the 7,000-token conversation they end with, because every call re-sends the history

Reading it: the bars are the input tokens billed on each call. They grow by the same amount every step, because each call re-sends the whole history. The line is the running total, and it curves upwards (quadratic growth). The dashed line is what you might naively expect: the size of the final conversation, billed once. The gap between the dashed line and the curve is why prompt caching, trimming tool outputs and keeping loops short are the main cost levers (primer.agents.cost, primer.agents.context).

In code: loop_input_tokens lists the tokens billed on each call of such a loop. estimate_tokens is the four-characters-per-token rule of thumb, and conversation_chars measures everything a call resends, which is how ScriptedLLM gives its fake replies realistic Usage.

Test yourself

5 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1What exactly happens when a model "uses a tool"?Think it through, then reveal

You send tool definitions (name, description, JSON Schema) with the messages. The model replies with a tool_use block and stop_reason: tool_use. Your code validates the arguments, runs the function, and sends a new request with the history plus a tool_result block carrying the same id. The model then answers or asks for another tool.

Question 2Why is the model never a security boundary for tool use?Think it through, then reveal

It only produces requests. Everything that actually happens goes through your code, so validation, permissions, approvals and rate limits all belong there, and a manipulated model can only ever ask.

Question 3The model asks for three tools in one reply. How do you send the results?Think it through, then reveal

Run them (concurrently if independent) and return all three tool_result blocks in a single user message, each with its own tool_use_id. For one that failed, return a tool_result with is_error: true and an actionable message rather than dropping it.

Question 4A 20-step agent loop costs far more than 20 times a single call. Why?Think it through, then reveal

Each call re-sends the full, growing history, so the input billed is a sum that grows with the square of the number of steps. Caching the stable prefix and trimming tool outputs attack exactly this.

Question 5Why test agents against a scripted model at all?Think it through, then reveal

Real models are non-deterministic and rarely misbehave on cue. A scripted model reproduces loops, bad arguments and injected instructions exactly, so the code that must handle them can be tested every time.

Primary sources

The papers behind this lesson

Schick et al., Toolformer: Language Models Can Teach Themselves to Use Tools (2023)

Showed a model can learn when to call external tools and how to use their results.

Read the annotated companion →The paper ↗
Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022)

The pattern of interleaving reasoning with tool calls that modern agent loops descend from.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Tool use overview: https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview
  • Implementing tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
  • Messages API reference: https://docs.claude.com/en/api/messages
  • Python SDK: https://github.com/anthropics/anthropic-sdk-python
  • Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.