rumblr Work in progressWIP

● The AI Primer · Lesson 51 · Part 2: building systems people rely on

Observability

recording every agent run so you can replay why it did what it did

You'll be able to explain Traces, OpenTelemetry GenAI attributes, the improvement loop

Members · open during launch 16 min7 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. A trace records one run as a tree of spans: the task, then each model call, tool call and retrieval, with inputs, outputs, tokens, latency, cost, and model and prompt versions.
  2. Use the OpenTelemetry GenAI attribute names so traces flow into existing monitoring or LLM tracing tools without custom glue.
  3. Debug from traces: find the first error span; the fix often belongs in a tool or in retrieval, not the prompt.
  4. Link traces to feedback and evals: production failure → trace → golden case → fix → CI → deploy.

Level 1

The practitioner's guide

In one sentence

Observability for an agent means recording every run as a tree of timed steps (the task, then each model call, tool call and retrieval, with what went in and what came out), so that "why did it do that?" is a lookup and not a guess.

When you need it

From the first day real people use the system. The tell: a user reports a wrong answer from yesterday and the only evidence is the answer itself. This lesson's worked example is exactly that case. A user asks "Where is order A-100?" (with a hyphen) and gets "I couldn't find order A-100." From the answer alone the order might not exist; the trace shows a tool error on one line, unknown order id A-100, and tells you the fix belongs in the tool that doesn't normalise hyphens, not in the prompt. The second tell is a bill you can't explain: in the lesson's six-order run, input tokens climb with every model call because each call resends the growing history, and only a per-call record makes that visible before the invoice does. You don't need a tracing platform for a script you run by hand, but even there, one structured record per model call (model, tokens, latency, cost) pays for itself the first time something is slow.

Your options

Six ways to see inside a run, from the cheapest to the most complete:

Option What it does What it guarantees What it costs Where it lives
Print statements and plain logs Lines of text as the agent runs You can read them, once; nothing ties one run's steps together Nothing up front; hours later, when you need to find one run among thousands Your code
Structured per-call records One record per model call with model, tokens, latency, cost, prompt version Dashboards for cost and volume; the bill becomes explainable A logging library and a few fields per call Your code
Traces with spans Each step is a timed span with attributes and a parent; a run is a tree under one trace id Replay any run; find the first error; see where the time goes Instrumenting each call (a with block), and somewhere to store the trees Your agent loop
A vendor's own SDK The tracing platform's client records the spans for you Fastest start Lock-in: changing vendors is a code change Your code, tied to one service
OpenTelemetry with the GenAI conventions The same spans in standard names (gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.tool.name) and the standard OTLP format Flows into the monitoring the organisation already runs and into any LLM tool; changing vendors is a collector setting Learning the conventions; running a collector Exporter and collector
An LLM tracing platform on top Traces, user feedback, prompt versions, datasets and evals in one place The loop from a thumbs-down to a golden test is a few clicks Hosting it yourself or sending traces (with user data in them) to a service A service

How to choose

Decide by what you'll need to answer, then by who else needs to read it.

  • You need to explain a bill: structured per-call records with tokens, model and cost; a sum over them is the invoice.
  • You need to explain a decision: traces. The first error span, or the first surprising output, shows where the chain broke. Nothing less lets you replay a run.
  • The organisation already runs Datadog, Grafana or Honeycomb: emit OpenTelemetry spans with the GenAI attribute names, and let the collector fan them out. Name your own attributes with a prefix (app.prompt.version) so they never collide with the standard's.
  • You want traces linked to feedback and evals: put an LLM tracing platform behind the collector; most of them accept OpenTelemetry, so the agent code doesn't change.
  • Whatever you pick, redact personal data before export, restrict who can read traces, set a retention limit, and keep tenants' traces apart. A trace holds the user's words.

What it costs

Instrumentation is cheap in code: a span is a with block around a call, and an exception inside it marks the span as an error by itself. It is not free in storage: a trace keeps every prompt and tool result, so this lesson's three-step run (two model calls, one tool call, 900 ms on its toy clock, 195 input and 38 output tokens, about a tenth of a cent) is a few kilobytes, and a million runs a day is gigabytes. Distributed tracing at scale has always answered that by sampling: Google's Dapper, the design OpenTelemetry descends from, kept overhead low by tracing a sample of requests through shared libraries rather than every request. Keep every trace while volume is small; sample when it isn't, but always keep the ones a user flagged or a monitor caught. Latency cost is negligible when export is asynchronous. The cost that matters is what you learn from the trace itself: in the lesson's six-order run, model calls take far more of the time than tool calls, so the latency wins are fewer model calls and smaller contexts, not faster tools.

What breaks

  • A log with no thread through it. Thousands of lines from thousands of runs, and no way to gather one run back together. A trace id on every span is the fix.
  • Fixing the prompt when the tool was wrong. Without the trace, the hyphen bug above looks like a model mistake. Read the first error span before editing anything.
  • Personal data in the trace store. Prompts and outputs hold names, emails and card numbers. Redact before export, and limit retention.
  • Vendor-shaped traces. Attributes named after one product's SDK don't flow anywhere else. Use the standard names and a collector.
  • Gaps between spans. Time outside any span (a queue, a retry wait, untraced code) is invisible in the tree. Instrument the retries and the guardrail checks too.
  • Traces nobody reads. Recording is half the job. Store user feedback against the trace id and route flagged traces to a person, so each one can become a golden case.

In the wild

OpenTelemetry is the open standard for traces, metrics and logs; its concepts page defines a trace as the path of a request through an application and a span as a unit of work with a name, a parent, start and end times, attributes, events and a status, which is exactly what this lesson's Span holds. Its GenAI semantic conventions, maintained in their own repository, standardise the span, metric and event names for model calls, tool calls, agents and MCP. Langfuse is an open-source, self-hostable platform that captures model and non-model calls, versions prompts, runs evaluations (LLM judge, code, human) and tracks cost per user, built on OpenTelemetry. Arize Phoenix is open source too, with tracing built on OpenTelemetry and OpenInference, evaluators and datasets for comparing versions. LangSmith and Braintrust play the same role as services. All of them descend from Dapper (Sigelman et al., 2010), which introduced the trace-and-span tree with propagated ids.

Go deeper

Level 2 builds the tracer in plain Python: a span, a with block that opens and closes it, a fake clock so timings are exact, the agent loop instrumented step by step, the export to OpenTelemetry-shaped records, the walk that finds the first error, and the path from a thumbs-down to a golden case. If you only needed to decide what to record and where to send it, you are done.

Level 2

How it works, from scratch

An aeroplane's flight recorder writes down every control input, every instrument reading and every radio call, each stamped with the time. When something goes wrong, investigators don't guess; they replay the flight.

A trace is a flight recorder for one agent run. It records the whole task and every step inside it (each model call, each tool call, each search) with its inputs, outputs, timing, token counts, cost, and which model and prompt version produced it. Without traces, "why did the agent refund that order?" is guesswork. With them, it's a lookup.

Chapter 1

Spans: the steps of a run, as a tree

Everyday picture A recipe card with sub-recipes: "make the lasagne" contains "make the sauce" and "make the béchamel", each with its own start and end time.

Each step is a span: a named, timed unit of work with key-value attributes (model name, tokens, tool name…), a status (ok or error), and a parent. The top span is the whole task; everything the agent does is a child of it. All spans of one run share a trace id, so they can be gathered back together.

Worked example "Where is order A100?" produces this trace, on a clock where a model call takes 400 ms and a tool call 100 ms:

invoke_agent support-agent  900 ms  [ok]
├─ chat scripted-large  400 ms  [ok]  in=68 out=24 tokens
├─ execute_tool lookup_order  100 ms  [ok]
└─ chat scripted-large  400 ms  [ok]  in=127 out=14 tokens

The model asked for a tool (first chat), the tool ran, and the model answered (second chat). Durations add up: 400 + 100 + 400 = 900 ms. The second model call has more input tokens because the conversation, including the tool result, is resent every step.

Figure 1 · Diagram

Reading it: the root box is the whole task; the three children are the steps, left to right in time. Every box carries its own timing and attributes, and every box can be opened to see exactly what went in and came out. Retrieval, sub-agents and guardrail checks become spans too, so the tree shows the whole path.

In code: Span holds one step: its name, kind, start and end times, attributes, status and children. walk visits a tree parents first, in time order, and render draws it as the text tree above.

Chapter 2

Recording spans as the agent runs

Figure 2 · Diagram

Reading it: the tracer sits beside the agent loop and is told when each step starts and ends; it never changes what the agent does. Opening a span before a call and closing it after is all instrumentation means. In Python it's a with tracer.span(...): block, and an exception inside the block marks the span as an error automatically.

In code: Tracer.span opens a child of whatever span is open, and closes and times it when the block ends; Span.fail marks an error that didn't raise. FakeClock makes the timings deterministic, and run_traced_agent is the agent loop in the diagram, recording every step.

Chapter 3

Speaking a common language: OpenTelemetry

Everyday picture Shipping containers: because every container has the same shape, any ship, crane or truck can carry any of them.

OpenTelemetry (OTel) is the open standard for traces, metrics and logs, and its GenAI semantic conventions standardize the attribute names for model and agent spans: gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.tool.name and so on. Use them, and your traces flow into whatever the organisation already runs (Datadog, Grafana, Honeycomb) or into LLM-specific tools (Langfuse, LangSmith, Arize Phoenix, Braintrust) with no custom glue. Attributes this module invents itself use an app. prefix, such as app.prompt.version and app.cost_usd.

Attribute Meaning Example
gen_ai.operation.name what kind of step chat, execute_tool, invoke_agent
gen_ai.request.model model asked for claude-opus-5
gen_ai.usage.input_tokens / output_tokens tokens billed 68 / 24
gen_ai.response.finish_reasons why the model stopped ["tool_use"]
gen_ai.tool.name, gen_ai.tool.call.id which tool, which call lookup_order, toolu_1
app.prompt.version which prompt produced this call support-v3

Figure 3 · Diagram

Reading it: the agent only knows how to hand spans to an exporter, which ships them in the standard OTLP format (OpenTelemetry Protocol). A collector fans them out to as many destinations as you like. Swapping vendors means changing the collector's configuration, not the agent. to_otel in this module produces OTLP-shaped span records.

Chapter 4

Debugging a failure from its trace

Everyday picture Replaying the flight recorder to the moment the warning light came on.

Worked example A user asks "Where is order A-100?" (with a hyphen). The agent's answer is an unhelpful "I couldn't find order A-100." From the answer alone the order might simply not exist. The trace shows why in one line:

invoke_agent support-agent  900 ms  [ok]
├─ chat scripted-large  400 ms  [ok]  in=68 out=24 tokens
├─ execute_tool lookup_order  100 ms  [error]  unknown order id A-100
└─ chat scripted-large  400 ms  [ok]  in=135 out=15 tokens

The model extracted the id exactly as typed and the tool doesn't normalize hyphens. first_error walks the tree and returns that span. The fix belongs in the tool (normalize ids, or return an error the model can act on, such as "ids look like A100"), not in the prompt, and it's the trace that tells you so.

Figure 4 · Drawn from the lesson's code

0 250 500 750 1000 1250 1500 1750 milliseconds since the task started Trace waterfall: model calls (blue) and tool calls (orange) invoke_agent support-agent chat model execute_tool lookup_order chat model execute_tool lookup_order chat model execute_tool lookup_order chat model

In a three-order run, four 400 ms model calls alternate back to back with three 100 ms tool calls, so model calls fill most of the 1,900 ms

Reading it: the classic trace view. Each row is a span, its bar placed on a shared time axis; the top row is the whole task. Model calls (blue) alternate with tool calls (orange), each one starting the moment the previous one ends: that staircase is the loop's rhythm, think, act, think, act. The blue bars are four times as long as the orange ones, so where the time goes is visible at a glance. In a real trace, a gap between two bars would be time spent outside any span (a queue, a retry wait, untraced code), and worth a look.

Figure 5 · Drawn from the lesson's code

1 2 3 4 5 6 7 model call number within one run 0 100 200 300 400 tokens Each call resends the growing conversation input output

Input tokens climb with every model call, because each call resends the growing history

Reading it: each bar is one model call in a run that looks up six orders one at a time. Input tokens climb with every step because each call resends the growing history. Traces make this visible per call, which is how you notice a verbose tool result or a runaway loop before the bill does.

Figure 6 · Drawn from the lesson's code

0 500 1000 1500 2000 2500 3000 total milliseconds in the six-order run model calls tool calls Where the time goes 2800 ms 600 ms

In the six-order run, model calls take far more of the time than tool calls do

Reading it: total milliseconds spent in model calls versus tool calls for that six-order run. Model calls dominate, so the biggest latency wins are fewer model calls (batch tool calls into one turn, run them in parallel) and smaller contexts, not faster tools.

In code: trace_totals adds up a trace's duration, model and tool calls, tokens, cost and errors, the numbers behind these figures.

Chapter 5

Closing the loop

Traces become most valuable when linked to feedback and evals. When a user clicks thumbs-down, the feedback is stored against the trace id; an engineer opens the exact trace, sees what went wrong, and turns it into a golden test case (golden_case_from_trace, see primer.agents.evals). The fix is made, the eval suite passes in CI (the automated checks run on every change), and it ships.

Figure 7 · Diagram

Reading it: follow the loop clockwise from production. Each lap turns a real failure into a permanent test, so quality ratchets up and the same bug can't come back unnoticed. Without traces the loop breaks at the second box: you know something failed but not why.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1What would you record for every agent run?Think it through, then reveal

A tree of spans: the task at the root, and each model call, tool call and retrieval as children, with inputs and outputs (redacted of personal data), token counts, latency, cost, model id, prompt version, tool arguments and results, errors, and the user and tenant it ran for. That's enough to replay any decision and to aggregate cost and quality.

Question 2A user says the agent gave a wrong answer yesterday. How do you find out why?Think it through, then reveal

Look up the trace by conversation or user id and walk it: did retrieval return the right document, did the model pick the right tool with the right arguments, did a tool error, did a guardrail fire? The first error span or the first surprising output shows where the chain broke, and the trace becomes a new golden test case.

Question 3Why use OpenTelemetry instead of a vendor's SDK directly?Think it through, then reveal

Portability and fit: the organisation's existing monitoring already speaks OTel, the GenAI conventions give every tool the same attribute names, and changing vendors becomes a collector configuration change rather than a code change.

Question 4What should you be careful about when logging prompts and outputs?Think it through, then reveal

They contain user data. Redact personal data before export, restrict who can read traces, set retention limits, and keep tenants' traces separated.

Primary sources

The papers behind this lesson

Sigelman et al., Dapper, a Large-Scale Distributed Systems Tracing Infrastructure (Google technical report, 2010), Introduced the trace-and-span tree model with propagated trace ids that OpenTelemetry, and therefore agent tracing, is built on.

The paper ↗

Researcher's shelf

Further reading

  • OpenTelemetry semantic conventions for generative AI: https://opentelemetry.io/docs/specs/semconv/gen-ai/
  • OpenTelemetry concepts: traces and spans: https://opentelemetry.io/docs/concepts/signals/traces/
  • Langfuse documentation: https://langfuse.com/docs
  • Arize Phoenix documentation: https://arize.com/docs/phoenix

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.