rumblr Work in progressWIP

● The AI Primer · Lesson 41 · Part 2: building systems people rely on

Tools

design, validation and safety

This lesson covers Tool design, validation, idempotency, approvals, least privilege

Members · open during launch 21 min9 figures and diagrams
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. The model writes a request (tool_use). Your code validates it, runs it, and returns a tool_result.
  2. Validate shape (JSON Schema) and meaning (business rules). Return every problem, each with its fix.
  3. Descriptions are prompts: say what the tool does, when to use it, and when not to.
  4. Prefer fewer, higher-level tools: punishes long chains of calls.
  5. Past a few dozen tools, load only the relevant ones per request.
  6. Writes get idempotency keys; irreversible or high-value actions get human approval; credentials get the narrowest scopes.

Level 2

How it works, from scratch

A language model can only produce text. A tool is how that text turns into an action: looking up a record, refunding a payment, sending an email. This lesson builds a tool registry from scratch and shows the six things that decide whether tools work in production: how a call really happens, validation, descriptions, granularity, how many tools you expose, and safety.

Chapter 1

What a tool call really is

Everyday picture You hire a new assistant who is brilliant but may not touch anything. When they need something done, they fill in a request form ("refund customer C-100, $20") and hand it to you. You decide whether to carry it out, and you hand back a note saying what happened. The assistant is the model. The form is a tool call. You are your code.

Tiny worked example Here's the whole exchange for "what's 2 + 3?" with one tool, written as the actual messages:

you -> model   tools: [{"name": "add", "description": "Add two integers.",
                        "input_schema": {"type": "object",
                          "properties": {"a": {"type": "integer"}, "b": {"type": "integer"}},
                          "required": ["a", "b"]}}]
               user: "what's 2 + 3?"
model -> you   tool_use {"id": "toolu_1", "name": "add", "input": {"a": 2, "b": 3}}
               stop_reason: "tool_use"          <- "please run this for me"
you            (validate the input, run add(2, 3) -> 5)
you -> model   user: tool_result {"tool_use_id": "toolu_1", "content": "5"}
model -> you   "2 + 3 = 5."   stop_reason: "end_turn"

input_schema is a JSON Schema, a JSON document that describes the shape other JSON must have: which fields exist, what type each is, and which are required. It's the "form template" the model fills in.

Figure 1 · Diagram

Reading it: follow the arrows top to bottom. The model's only arrow towards the real system goes through your code. Nothing the model writes can touch the real system unless your code decides to act on it. That middle box, "check the form", is where the rest of this lesson lives.

The code ToolRegistry.register stores a Python function with its name, description and schema; ToolRegistry.definitions() produces the list you send to the model; ToolRegistry.call() runs a call through every check.

Why it matters The separation is the security model. A model can be wrong or tricked. Your code is where you enforce what's allowed.

Chapter 2

Validate every call before running it

Everyday picture A bank clerk checks a withdrawal slip twice. Is it filled in properly: every box, a real date, a positive amount? Then does it make sense: does this account exist, does it have the money? Only then do they open the drawer.

Tiny worked example A payment tool takes an amount, a currency (EUR, GBP or USD) and an optional pay_on date. The model sends {"ammount": 5, "currency": "usd", "pay_on": "next friday"}. The validator returns every problem at once, each one saying what a correct value looks like:

$: missing required field 'amount'
$: unexpected field 'ammount' (allowed: amount, currency, pay_on)
$.currency: must be one of ['EUR', 'GBP', 'USD'], got 'usd'
$.pay_on: 'next friday' does not match the required format. Expected: date as YYYY-MM-DD

Compare that to invalid input. With the first, the model fixes all four mistakes in one retry. With the second, it guesses.

Figure 2 · Diagram

Reading it: a call enters at the top and must pass every diamond to reach run the tool at the bottom. Each exit on the side is a named status in ToolOutcome, and each returns text the model can act on. The order is deliberate: permissions come before anything that reveals how the tool works, and cheap checks come before expensive ones. Approval and idempotency sit last, right next to the action they protect.

Why it matters "Structured output" and strict mode guarantee the shape of the arguments, not that they're true. "C-999" matches the pattern ^C-\d+$ perfectly and is still a customer that doesn't exist. Schema checks and business-rule checks are both required.

With the Claude API you can add "strict": true to a tool definition (definitions(strict=True) here). The API then guarantees the model's arguments validate against the schema. You still need the business-rule checks.

In code: validate checks a value against a JSON Schema subset and returns every problem with its path. ToolRegistry.call walks the diamonds above in order and wraps the result in a ToolOutcome, whose ToolOutcome.is_error says whether the model should treat it as a failure.

Chapter 3

Descriptions are prompts

Everyday picture A wall of drawers labelled "stuff", "things" and "items". Even a careful person opens the wrong one. Relabel them "Policies (PTO, travel, VPN). Not customers", and nobody hesitates.

Tiny worked example A stand-in model (pick_tool) chooses the tool whose name and description share the most words with the request. For "billing status for customer 1042", the vague descriptions share zero words with every tool (a three-way tie, so it guesses the first). The precise lookup_record description shares "billing", "status" and "customer" and wins.

Figure 3 · Chart

vague descriptions precise descriptions 0 1 2 3 4 5 6 requests routed to the right tool Same model, same requests, new descriptions 2/6 6/6

Vague descriptions pick the right tool for 2 of 6 requests; precise descriptions of the same tools get all 6 right

Reading it: two bars, one per set of descriptions, each out of the same six labelled requests. With vague descriptions the model gets 2 of 6, and only the two that happen to belong to the first tool, which it picks every time it can't tell them apart. With descriptions that say what each tool is for, when to use it and when not to, it gets 6 of 6. Only the descriptions changed.

Why it matters Real models are far better readers than word overlap, but they choose from exactly the same text. Precise descriptions, a few parameters, enums instead of free text, and error messages that explain the fix remove a large share of agent errors.

In code: selection_accuracy runs pick_tool over the labelled requests and counts the correct picks, which is what the figure plots for VAGUE_TOOLS and PRECISE_TOOLS.

Chapter 4

Fewer, higher-level tools

Everyday picture Asking an assistant to "book my trip to Denver" versus dictating five separate forms (find flight, hold seat, find hotel, reserve room, add to calendar). Every hand-off is another chance to drop something.

Figure 4 · Diagram

Reading it: on the left the model has to pick five tools in the right order and carry each output into the next input by hand. On the right the same work is one call, and the sequencing lives in ordinary tested code.

The chance that a chain of calls all succeed:

Level 3: the formula and its symbols

Symbols

Symbol Meaning
probability that one call is chosen and filled in correctly
number of calls in the chain

In words: multiply the per-call success rate by itself once per call.

On the example: with and , , so about 1 run in 7 fails. With one high-level call it's .

In Python:

p = 0.97
# five calls that must all succeed
round(p ** 5, 3)  # → 0.859
# about 1 run in 7 fails
round(1 - p ** 5, 2)  # → 0.14
# one high-level call
p ** 1  # → 0.97

Figure 5 · Chart

2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 calls in the chain (n) 0.0 0.2 0.4 0.6 0.8 1.0 P(every call succeeds) = p^n Reliability compounds: fewer, higher-level tools p = 0.90 per call p = 0.95 per call p = 0.97 per call p = 0.99 per call

Success falls with every extra call: at 99% per call twenty calls succeed about 82% of the time, at 90% ten calls barely a third

Reading it: the x-axis is how many calls the job takes, and the y-axis is the chance the whole job succeeds. Each curve is a different per-call reliability. Even at 99% per call, twenty calls succeed only about 82% of the time, and at 90% per call, ten calls succeed barely a third of the time. Fewer calls is the cheapest reliability you can buy.

In code: chain_success evaluates .

Chapter 5

Too many tools: load only the relevant ones

Everyday picture A warehouse versus a toolbox. You don't send a plumber into a warehouse of 500 tools; you hand them the five they need for today's job.

Tiny worked example A catalogue of 30 tools. For "I forgot my password and I'm locked out", select_tools embeds the request, compares it to every tool description, and sends the model only the top 5, with reset_password first.

Figure 6 · Diagram

Reading it: this is retrieval, the same machinery as RAG, but the "documents" are tool descriptions. The catalogue is embedded ahead of time; per request you embed one string and take the top matches.

Figure 7 · Chart

reset_password enroll_mfa check_vpn_status renew_device_certificate request_laptop report_phishing fix_printer submit_expense book_travel get_per_diem request_pto get_pto_balance report_sick_day request_parental_leave get_payslip get_salary_band start_onboarding list_invoices list_payments reconcile_invoices create_invoice send_email search_policies open_ticket close_ticket get_org_chart book_meeting_room order_office_supplies translate_text summarize_document I forgot my password and I'm locked out submit my car mileage expense how many vacation days are left suspicious email asking for my login reconcile vendor invoices for Q3 Cosine similarity: each request lights up a few tools −0.2 0.0 0.2 0.4 0.6 0.8

Each request is similar to only a small cluster of related tools and dim against the rest of the catalogue

Reading it: each row is a user request and each column a tool from the catalogue. Brighter cells mean more similar. Each request lights up a small cluster of related tools and stays dark everywhere else. Sending only that cluster keeps the model's choice small, which is why accuracy holds up as the catalogue grows.

Why it matters Selection accuracy drops as the tool list grows past a few dozen, and every definition costs input tokens on every call. The alternatives are routing the request to a sub-agent that holds only the relevant tools, or using a provider's built-in tool search.

Chapter 6

Safety: idempotency, dry runs, approval, least privilege

Idempotency. Everyday picture: pressing a lift button twice still takes you to the floor once. An operation is idempotent when doing it twice has the same effect as doing it once. Worked example: a refund call times out after the bank processed it, so the agent retries.

Figure 8 · Diagram

Reading it: the first reply is lost on the way back (the crossed arrow), so the agent can't know the refund happened and sensibly retries. Because the retry carries the same key, the registry returns the stored result instead of refunding again. Without the key, the customer gets $40.

Dry run. A rehearsal: run every check, then describe the action instead of doing it. Useful for previews and for shadow-mode rollouts (primer.agents.deployment).

Human approval. Everyday picture: a manager signs off on spending over a limit. Irreversible actions, or ones over a threshold such as refunds over $500, park as awaiting_approval until a person decides.

Figure 9 · Diagram

Reading it: the approval gate sits between the model's request and the action. Both branches return text the model can act on. "Do not retry" in the declined message matters, because otherwise a persistent model asks again.

Least privilege. Everyday picture: a valet key starts the car but won't open the boot. Give each agent credentials (scopes, named permissions such as payments:write) for its job only. An agent that reads tickets holds tickets:read, so even if a malicious email tricks it, it cannot issue refunds.

In code: a Tool carries its safety settings: the scopes it needs, whether it reads, writes or acts irreversibly, and an optional approval rule, which Tool.requires_approval combines. ToolRegistry.call enforces them, given the caller's credentials, an idempotency key, a dry-run flag and an approver.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: How do you design tools so the model picks the right one with correct arguments?Think it through, then reveal

A: Give each tool one clear job and a description that says what it does, when to use it and when not to. Keep parameters few, use enums for closed sets, and put the format in the schema description ("date as YYYY-MM-DD"). Prefer one high-level tool over several low-level ones. Validate every call and return actionable errors so the model can self-correct. Measure tool selection accuracy in evals, and load tools dynamically when there are many.

Question 2Q: The schema says customer_id must match ^C-\d+$. Is that enough validation?Think it through, then reveal

A: No. The schema checks shape, not truth. C-999 is well-formed and may not exist. Add a business-rule check before acting, and return an error that tells the model how to find the right id.

Question 3Q: A refund call timed out. Should the agent retry?Think it through, then reveal

A: Only if the call is idempotent. Send an idempotency key with every write, and the retry returns the original result instead of refunding twice.

Question 4Q: Why not give the agent one admin credential for everything?Think it through, then reveal

A: Blast radius. If the agent is tricked (for example by prompt injection in a document it reads), it can do anything the credential allows. Scope credentials per tool and per agent, and put irreversible actions behind human approval.

Primary sources

The papers behind this lesson

Schick et al., Toolformer: Language Models Can Teach Themselves to Use Tools (2023).

It showed a language model can learn when to call an API, which one and with what arguments, by keeping only the self-generated calls that made its predictions better, which is the idea behind tool calling being trained into today's models.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Anthropic, Writing effective tools for agents: https://www.anthropic.com/engineering/writing-tools-for-agents
  • Anthropic, Building effective agents (appendix on prompt-engineering tools): https://www.anthropic.com/engineering/building-effective-agents
  • Claude tool use overview: https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview
  • Implementing tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
  • Understanding JSON Schema: https://json-schema.org/understanding-json-schema
  • Stripe, idempotent requests (the canonical explanation): https://docs.stripe.com/api/idempotent_requests

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.