At a glance
Key takeaways
- A trace records one run as a tree of spans: the task, then each model call, tool call and retrieval, with inputs, outputs, tokens, latency, cost, and model and prompt versions.
- Use the OpenTelemetry GenAI attribute names so traces flow into existing monitoring or LLM tracing tools without custom glue.
- Debug from traces: find the first error span; the fix often belongs in a tool or in retrieval, not the prompt.
- Link traces to feedback and evals: production failure → trace → golden case → fix → CI → deploy.
Level 2
How it works, from scratch
An aeroplane's flight recorder writes down every control input, every instrument reading and every radio call, each stamped with the time. When something goes wrong, investigators don't guess; they replay the flight.
A trace is a flight recorder for one agent run. It records the whole task and every step inside it (each model call, each tool call, each search) with its inputs, outputs, timing, token counts, cost, and which model and prompt version produced it. Without traces, "why did the agent refund that order?" is guesswork. With them, it's a lookup.
Chapter 1
Spans: the steps of a run, as a tree
Everyday picture A recipe card with sub-recipes: "make the lasagne" contains "make the sauce" and "make the béchamel", each with its own start and end time.
Each step is a span: a named, timed unit of work with key-value attributes (model name, tokens, tool name…), a status (ok or error), and a parent. The top span is the whole task; everything the agent does is a child of it. All spans of one run share a trace id, so they can be gathered back together.
Worked example "Where is order A100?" produces this trace, on a clock where a model call takes 400 ms and a tool call 100 ms:
invoke_agent support-agent 900 ms [ok]
├─ chat scripted-large 400 ms [ok] in=68 out=24 tokens
├─ execute_tool lookup_order 100 ms [ok]
└─ chat scripted-large 400 ms [ok] in=127 out=14 tokens
The model asked for a tool (first chat), the tool ran, and the model
answered (second chat). Durations add up: 400 + 100 + 400 = 900 ms. The
second model call has more input tokens because the conversation, including
the tool result, is resent every step.
Figure 1 · Diagram
flowchart TD R["invoke_agent support-agent<br/>trace id tr-0001, 900 ms"] --> C1["chat: model call 1<br/>400 ms, 68 in / 24 out"] R --> T1["execute_tool lookup_order<br/>100 ms, args: A100"] R --> C2["chat: model call 2<br/>400 ms, 127 in / 14 out"]
In code: Span holds one step: its name, kind, start and end times,
attributes, status and children. walk visits a tree parents first, in time
order, and render draws it as the text tree above.
Chapter 2
Recording spans as the agent runs
Figure 2 · Diagram
sequenceDiagram participant U as User participant A as Agent loop participant T as Tracer participant M as Model participant K as Tool U->>A: Where is order A100? A->>T: open span invoke_agent A->>T: open span chat A->>M: messages + tools M-->>A: tool_use lookup_order(A100) A->>T: close span chat (tokens, model, finish reason) A->>T: open span execute_tool A->>K: lookup_order(A100) K-->>A: shipped A->>T: close span execute_tool A->>T: open span chat A->>M: messages + tool result M-->>A: "Order A100 has shipped." A->>T: close span chat A->>T: close span invoke_agent A-->>U: Order A100 has shipped.
with tracer.span(...): block, and an exception inside the
block marks the span as an error automatically.In code: Tracer.span opens a child of whatever span is open, and
closes and times it when the block ends; Span.fail marks an error that
didn't raise. FakeClock makes the timings deterministic, and
run_traced_agent is the agent loop in the diagram, recording every step.
Chapter 3
Speaking a common language: OpenTelemetry
Everyday picture Shipping containers: because every container has the same shape, any ship, crane or truck can carry any of them.
OpenTelemetry (OTel) is the open standard for traces, metrics and logs,
and its GenAI semantic conventions standardize the attribute names for
model and agent spans: gen_ai.request.model, gen_ai.usage.input_tokens,
gen_ai.tool.name and so on. Use them, and your traces flow into whatever
the organisation already runs (Datadog, Grafana, Honeycomb) or into
LLM-specific tools (Langfuse, LangSmith, Arize Phoenix, Braintrust) with no
custom glue. Attributes this module invents itself use an app. prefix,
such as app.prompt.version and app.cost_usd.
| Attribute | Meaning | Example |
|---|---|---|
gen_ai.operation.name |
what kind of step | chat, execute_tool, invoke_agent |
gen_ai.request.model |
model asked for | claude-opus-5 |
gen_ai.usage.input_tokens / output_tokens |
tokens billed | 68 / 24 |
gen_ai.response.finish_reasons |
why the model stopped | ["tool_use"] |
gen_ai.tool.name, gen_ai.tool.call.id |
which tool, which call | lookup_order, toolu_1 |
app.prompt.version |
which prompt produced this call | support-v3 |
Figure 3 · Diagram
flowchart LR A[Agent code<br/>+ tracer] -->|spans| X[Exporter] X -->|OTLP| C[Collector] C --> D[Existing monitoring<br/>Datadog, Grafana] C --> L[LLM tracing tools<br/>Langfuse, Phoenix] C --> W[Warehouse<br/>for evals and audits]
to_otel in this module produces OTLP-shaped span records.Chapter 4
Debugging a failure from its trace
Everyday picture Replaying the flight recorder to the moment the warning light came on.
Worked example A user asks "Where is order A-100?" (with a hyphen). The agent's answer is an unhelpful "I couldn't find order A-100." From the answer alone the order might simply not exist. The trace shows why in one line:
invoke_agent support-agent 900 ms [ok]
├─ chat scripted-large 400 ms [ok] in=68 out=24 tokens
├─ execute_tool lookup_order 100 ms [error] unknown order id A-100
└─ chat scripted-large 400 ms [ok] in=135 out=15 tokens
The model extracted the id exactly as typed and the tool doesn't normalize
hyphens. first_error walks the tree and returns that span. The fix
belongs in the tool (normalize ids, or return an error the model can act
on, such as "ids look like A100"), not in the prompt, and it's the trace
that tells you so.
Figure 4 · Chart
In a three-order run, four 400 ms model calls alternate back to back with three 100 ms tool calls, so model calls fill most of the 1,900 ms
Figure 5 · Chart
Input tokens climb with every model call, because each call resends the growing history
Figure 6 · Chart
In the six-order run, model calls take far more of the time than tool calls do
In code: trace_totals adds up a trace's duration, model and tool calls,
tokens, cost and errors, the numbers behind these figures.
Chapter 5
Closing the loop
Traces become most valuable when linked to feedback and evals. When a user
clicks thumbs-down, the feedback is stored against the trace id; an engineer
opens the exact trace, sees what went wrong, and turns it into a golden
test case (golden_case_from_trace, see primer.agents.evals). The fix is
made, the eval suite passes in CI (the automated checks run on every
change), and it ships.
Figure 7 · Diagram
flowchart LR P[Production traffic] --> T[Traces] T --> F[Failures flagged<br/>by users or monitors] F --> G[Add to golden set] G --> X[Fix prompt, tools<br/>or retrieval] X --> E[Evals pass in CI] E --> D[Deploy] D --> P
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1What would you record for every agent run?Think it through, then reveal
A tree of spans: the task at the root, and each model call, tool call and retrieval as children, with inputs and outputs (redacted of personal data), token counts, latency, cost, model id, prompt version, tool arguments and results, errors, and the user and tenant it ran for. That's enough to replay any decision and to aggregate cost and quality.
Question 2A user says the agent gave a wrong answer yesterday. How do you find out why?Think it through, then reveal
Look up the trace by conversation or user id and walk it: did retrieval return the right document, did the model pick the right tool with the right arguments, did a tool error, did a guardrail fire? The first error span or the first surprising output shows where the chain broke, and the trace becomes a new golden test case.
Question 3Why use OpenTelemetry instead of a vendor's SDK directly?Think it through, then reveal
Portability and fit: the organisation's existing monitoring already speaks OTel, the GenAI conventions give every tool the same attribute names, and changing vendors becomes a collector configuration change rather than a code change.
Question 4What should you be careful about when logging prompts and outputs?Think it through, then reveal
They contain user data. Redact personal data before export, restrict who can read traces, set retention limits, and keep tenants' traces separated.
Primary sources
The papers behind this lesson
Sigelman et al., Dapper, a Large-Scale Distributed Systems Tracing Infrastructure (Google technical report, 2010), Introduced the trace-and-span tree model with propagated trace ids that OpenTelemetry, and therefore agent tracing, is built on.
The paper ↗Researcher's shelf
Further reading
- OpenTelemetry semantic conventions for generative AI: https://opentelemetry.io/docs/specs/semconv/gen-ai/
- OpenTelemetry concepts: traces and spans: https://opentelemetry.io/docs/concepts/signals/traces/
- Langfuse documentation: https://langfuse.com/docs
- Arize Phoenix documentation: https://arize.com/docs/phoenix
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.