rumblr Work in progressWIP

● The AI Primer · Lesson 51 · Part 2: building systems people rely on

Observability

recording every agent run so you can replay why it did what it did

This lesson covers Traces, OpenTelemetry GenAI attributes, the improvement loop

Members · open during launch 16 min7 figures and diagrams
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. A trace records one run as a tree of spans: the task, then each model call, tool call and retrieval, with inputs, outputs, tokens, latency, cost, and model and prompt versions.
  2. Use the OpenTelemetry GenAI attribute names so traces flow into existing monitoring or LLM tracing tools without custom glue.
  3. Debug from traces: find the first error span; the fix often belongs in a tool or in retrieval, not the prompt.
  4. Link traces to feedback and evals: production failure → trace → golden case → fix → CI → deploy.

Level 2

How it works, from scratch

An aeroplane's flight recorder writes down every control input, every instrument reading and every radio call, each stamped with the time. When something goes wrong, investigators don't guess; they replay the flight.

A trace is a flight recorder for one agent run. It records the whole task and every step inside it (each model call, each tool call, each search) with its inputs, outputs, timing, token counts, cost, and which model and prompt version produced it. Without traces, "why did the agent refund that order?" is guesswork. With them, it's a lookup.

Chapter 1

Spans: the steps of a run, as a tree

Everyday picture A recipe card with sub-recipes: "make the lasagne" contains "make the sauce" and "make the béchamel", each with its own start and end time.

Each step is a span: a named, timed unit of work with key-value attributes (model name, tokens, tool name…), a status (ok or error), and a parent. The top span is the whole task; everything the agent does is a child of it. All spans of one run share a trace id, so they can be gathered back together.

Worked example "Where is order A100?" produces this trace, on a clock where a model call takes 400 ms and a tool call 100 ms:

invoke_agent support-agent  900 ms  [ok]
├─ chat scripted-large  400 ms  [ok]  in=68 out=24 tokens
├─ execute_tool lookup_order  100 ms  [ok]
└─ chat scripted-large  400 ms  [ok]  in=127 out=14 tokens

The model asked for a tool (first chat), the tool ran, and the model answered (second chat). Durations add up: 400 + 100 + 400 = 900 ms. The second model call has more input tokens because the conversation, including the tool result, is resent every step.

Figure 1 · Diagram

Reading it: the root box is the whole task; the three children are the steps, left to right in time. Every box carries its own timing and attributes, and every box can be opened to see exactly what went in and came out. Retrieval, sub-agents and guardrail checks become spans too, so the tree shows the whole path.

In code: Span holds one step: its name, kind, start and end times, attributes, status and children. walk visits a tree parents first, in time order, and render draws it as the text tree above.

Chapter 2

Recording spans as the agent runs

Figure 2 · Diagram

Reading it: the tracer sits beside the agent loop and is told when each step starts and ends; it never changes what the agent does. Opening a span before a call and closing it after is all instrumentation means. In Python it's a with tracer.span(...): block, and an exception inside the block marks the span as an error automatically.

In code: Tracer.span opens a child of whatever span is open, and closes and times it when the block ends; Span.fail marks an error that didn't raise. FakeClock makes the timings deterministic, and run_traced_agent is the agent loop in the diagram, recording every step.

Chapter 3

Speaking a common language: OpenTelemetry

Everyday picture Shipping containers: because every container has the same shape, any ship, crane or truck can carry any of them.

OpenTelemetry (OTel) is the open standard for traces, metrics and logs, and its GenAI semantic conventions standardize the attribute names for model and agent spans: gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.tool.name and so on. Use them, and your traces flow into whatever the organisation already runs (Datadog, Grafana, Honeycomb) or into LLM-specific tools (Langfuse, LangSmith, Arize Phoenix, Braintrust) with no custom glue. Attributes this module invents itself use an app. prefix, such as app.prompt.version and app.cost_usd.

Attribute Meaning Example
gen_ai.operation.name what kind of step chat, execute_tool, invoke_agent
gen_ai.request.model model asked for claude-opus-5
gen_ai.usage.input_tokens / output_tokens tokens billed 68 / 24
gen_ai.response.finish_reasons why the model stopped ["tool_use"]
gen_ai.tool.name, gen_ai.tool.call.id which tool, which call lookup_order, toolu_1
app.prompt.version which prompt produced this call support-v3

Figure 3 · Diagram

Reading it: the agent only knows how to hand spans to an exporter, which ships them in the standard OTLP format (OpenTelemetry Protocol). A collector fans them out to as many destinations as you like. Swapping vendors means changing the collector's configuration, not the agent. to_otel in this module produces OTLP-shaped span records.

Chapter 4

Debugging a failure from its trace

Everyday picture Replaying the flight recorder to the moment the warning light came on.

Worked example A user asks "Where is order A-100?" (with a hyphen). The agent's answer is an unhelpful "I couldn't find order A-100." From the answer alone the order might simply not exist. The trace shows why in one line:

invoke_agent support-agent  900 ms  [ok]
├─ chat scripted-large  400 ms  [ok]  in=68 out=24 tokens
├─ execute_tool lookup_order  100 ms  [error]  unknown order id A-100
└─ chat scripted-large  400 ms  [ok]  in=135 out=15 tokens

The model extracted the id exactly as typed and the tool doesn't normalize hyphens. first_error walks the tree and returns that span. The fix belongs in the tool (normalize ids, or return an error the model can act on, such as "ids look like A100"), not in the prompt, and it's the trace that tells you so.

Figure 4 · Chart

0 250 500 750 1000 1250 1500 1750 milliseconds since the task started Trace waterfall: model calls (blue) and tool calls (orange) invoke_agent support-agent chat model execute_tool lookup_order chat model execute_tool lookup_order chat model execute_tool lookup_order chat model

In a three-order run, four 400 ms model calls alternate back to back with three 100 ms tool calls, so model calls fill most of the 1,900 ms

Reading it: the classic trace view. Each row is a span, its bar placed on a shared time axis; the top row is the whole task. Model calls (blue) alternate with tool calls (orange), each one starting the moment the previous one ends: that staircase is the loop's rhythm, think, act, think, act. The blue bars are four times as long as the orange ones, so where the time goes is visible at a glance. In a real trace, a gap between two bars would be time spent outside any span (a queue, a retry wait, untraced code), and worth a look.

Figure 5 · Chart

1 2 3 4 5 6 7 model call number within one run 0 100 200 300 400 tokens Each call resends the growing conversation input output

Input tokens climb with every model call, because each call resends the growing history

Reading it: each bar is one model call in a run that looks up six orders one at a time. Input tokens climb with every step because each call resends the growing history. Traces make this visible per call, which is how you notice a verbose tool result or a runaway loop before the bill does.

Figure 6 · Chart

0 500 1000 1500 2000 2500 3000 total milliseconds in the six-order run model calls tool calls Where the time goes 2800 ms 600 ms

In the six-order run, model calls take far more of the time than tool calls do

Reading it: total milliseconds spent in model calls versus tool calls for that six-order run. Model calls dominate, so the biggest latency wins are fewer model calls (batch tool calls into one turn, run them in parallel) and smaller contexts, not faster tools.

In code: trace_totals adds up a trace's duration, model and tool calls, tokens, cost and errors, the numbers behind these figures.

Chapter 5

Closing the loop

Traces become most valuable when linked to feedback and evals. When a user clicks thumbs-down, the feedback is stored against the trace id; an engineer opens the exact trace, sees what went wrong, and turns it into a golden test case (golden_case_from_trace, see primer.agents.evals). The fix is made, the eval suite passes in CI (the automated checks run on every change), and it ships.

Figure 7 · Diagram

Reading it: follow the loop clockwise from production. Each lap turns a real failure into a permanent test, so quality ratchets up and the same bug can't come back unnoticed. Without traces the loop breaks at the second box: you know something failed but not why.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1What would you record for every agent run?Think it through, then reveal

A tree of spans: the task at the root, and each model call, tool call and retrieval as children, with inputs and outputs (redacted of personal data), token counts, latency, cost, model id, prompt version, tool arguments and results, errors, and the user and tenant it ran for. That's enough to replay any decision and to aggregate cost and quality.

Question 2A user says the agent gave a wrong answer yesterday. How do you find out why?Think it through, then reveal

Look up the trace by conversation or user id and walk it: did retrieval return the right document, did the model pick the right tool with the right arguments, did a tool error, did a guardrail fire? The first error span or the first surprising output shows where the chain broke, and the trace becomes a new golden test case.

Question 3Why use OpenTelemetry instead of a vendor's SDK directly?Think it through, then reveal

Portability and fit: the organisation's existing monitoring already speaks OTel, the GenAI conventions give every tool the same attribute names, and changing vendors becomes a collector configuration change rather than a code change.

Question 4What should you be careful about when logging prompts and outputs?Think it through, then reveal

They contain user data. Redact personal data before export, restrict who can read traces, set retention limits, and keep tenants' traces separated.

Primary sources

The papers behind this lesson

Sigelman et al., Dapper, a Large-Scale Distributed Systems Tracing Infrastructure (Google technical report, 2010), Introduced the trace-and-span tree model with propagated trace ids that OpenTelemetry, and therefore agent tracing, is built on.

The paper ↗

Researcher's shelf

Further reading

  • OpenTelemetry semantic conventions for generative AI: https://opentelemetry.io/docs/specs/semconv/gen-ai/
  • OpenTelemetry concepts: traces and spans: https://opentelemetry.io/docs/concepts/signals/traces/
  • Langfuse documentation: https://langfuse.com/docs
  • Arize Phoenix documentation: https://arize.com/docs/phoenix

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.