At a glance
Key takeaways
- Earn autonomy in steps: shadow mode, then human approval of each action, then autonomy for low-risk actions only; demote on any incident or drift.
- Treat prompts as code: versioned (content hashes), reviewed, evaluated, and one step from rollback.
- Release to a small canary share first and grow it only while its metrics match the current version; roll back automatically otherwise.
- Bound the blast radius: kill switches per tenant and globally, and rate limits on actions.
- Keep a tamper-evident audit log (a hash chain) linking every action to its user, authority and trace.
Level 2
How it works, from scratch
A new pilot doesn't get the captain's seat on day one. First they sit in the right seat and call out every move they would make while the captain flies. Then they fly with the captain's hands hovering over the controls. Then they fly routine legs alone, while the captain still handles storms. Trust is earned in steps, with evidence at each one, and it can be taken back.
Deploying an agent that takes real actions (refunds, emails, changes to a company's records) works the same way. This lesson builds each mechanism: graduated autonomy, prompt versioning, canary releases, kill switches, rate limits, and a tamper-evident audit log.
Figure 1 · Diagram
flowchart LR
P[Proposed action] --> K{Kill switch<br/>on?}
K -->|yes| X[Blocked]
K -->|no| R{Rate limit<br/>allows it?}
R -->|no| X
R -->|yes| A{Autonomy level<br/>and risk allow<br/>acting alone?}
A -->|no| H[Queue for<br/>human approval]
A -->|yes| E[Execute]
H -->|approved| E
E --> L[Append to<br/>audit log]
In code: each gate is one call: KillSwitch.allowed, then
TokenBucket.allow, then AutonomyController.may_act_alone, and finally
AuditLog.append for whatever executes.
Chapter 1
Graduated autonomy: shadow mode, approval, then autonomy
Everyday picture The trainee pilot again: call out moves (shadow), fly with the captain ready to take over (approval), fly routine legs alone (autonomy for low-risk actions), and hand back the controls after any incident.
In shadow mode the agent runs on real inputs and records what it would do, but takes no action; people keep doing the work. You compare its decisions with theirs. When agreement is high enough, it proposes actions and a person approves each one. When approvals are nearly always "yes", it may act alone on low-risk actions, while high-risk ones keep needing a person. Any incident or drop in quality sends it back a level.
Worked example Ten support tickets in shadow mode:
| Agent proposed | Human did | Agree? |
|---|---|---|
| refund | refund | yes |
| refund | escalate | no |
| reply ×5 | reply ×5 | yes ×5 |
| refund | refund | yes |
| refund | escalate | no |
| refund | refund | yes |
Overall agreement is 8/10 = 80%. Broken down by what the agent proposed, replies agree 5/5 = 100% but refunds only 3/5 = 60%. The breakdown tells you what to let the agent do first: replies, not refunds.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| number of shadow decisions compared | 10 | |
| what the agent proposed for case | refund | |
| what the human actually did for case | escalate | |
| 1 if the condition is true, 0 if not (an indicator) | 0 for case 2 | |
| add up over all cases |
In words: agreement is the share of cases where the agent proposed exactly what the human did.
On the worked example: eight of the ten indicators are 1, so agreement = 8/10 = 0.8.
Level 3: in Python
a = ["refund", "refund"] + ["reply"] * 5 + ["refund", "refund", "refund"]
h = ["refund", "escalate"] + ["reply"] * 5 + ["refund", "escalate", "refund"]
# 𝟙[a_i = h_i]
indicators = [1 if a_i == h_i else 0 for a_i, h_i in zip(a, h)]
indicators # → [1, 0, 1, 1, 1, 1, 1, 1, 0, 1]
# (1/N) Σ over i
sum(indicators) / len(indicators) # → 0.8
Figure 2 · Diagram
stateDiagram-v2 state "Shadow mode: record, don't act" as Shadow state "Human approves each action" as Approve state "Autonomous for low-risk actions" as Auto [*] --> Shadow Shadow --> Approve: agreement ≥ 95% over the last 50 decisions Approve --> Auto: ≥ 98% of the last 50 approved, no incidents Auto --> Approve: incident or quality drift Approve --> Shadow: incident
Figure 3 · Chart
Rolling agreement with humans must first clear the 95% bar over a full 50-decision window, and only then is the agent promoted
In code: shadow_agreement computes agreement overall and per proposed
action. AutonomyController is the state diagram: AutonomyController.record_shadow
and AutonomyController.record_approval promote once a full window clears
the bar, while AutonomyController.record_incident and
AutonomyController.record_quality demote at once.
Chapter 2
Prompts are code
Everyday picture A restaurant's recipe binder where every recipe change gets a new revision number, the old version stays in the binder, and the kitchen can switch back in one step if customers complain.
A change to a prompt can break behaviour as badly as a code change, so treat
it the same way: version it, review it, run the evals before release
(primer.agents.evals), and keep the previous version one step away.
PromptRegistry is content-addressed: a version's name includes a
hash of its text. A hash (here SHA-256) is a fixed-length fingerprint:
the same text always gives the same fingerprint, and any change, even one
character, gives a completely different one.
Worked example "Be concise." hashes to ab018bdb…, so its version is
support@ab018bdb. Registering the same text again returns the same
version; "Be concise and friendly." gets a new one. Every trace records
which version produced each model call (primer.agents.observability), so
a behaviour change can be tied to the exact prompt edit that caused it.
In code: PromptRegistry.register hashes the text into a version name
and makes it active; PromptRegistry.rollback steps back one version;
PromptRegistry.active_text returns the prompt currently in use.
Chapter 3
Canary releases: try it on a few users first
Everyday picture Miners once carried a canary into the mine: if the air turned bad, the canary showed it before the miners were harmed. A canary release sends a small share of traffic to the new version, compares its metrics with the current version (the control), and grows the share only while the canary stays healthy.
Worked example Steps 1% → 5% → 25% → 50% → 100%, allowed drop 2 points, at least 100 canary tasks before judging:
| Control success | Canary success | Canary tasks | Decision |
|---|---|---|---|
| 95% | 95% | 100 | advance from 1% to 5% |
| 95% | 90% | 100 | 90 < 95 − 2: roll back to 0% |
| 95% | 90% | 10 | too few samples: stay at 1% |
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| measured success rate on canary traffic (the hat means "measured from samples") | 0.90 | |
| measured success rate on the current version | 0.95 | |
| the largest drop you'll tolerate (Greek delta) | 0.02 |
In words: roll back when the new version's success rate falls more than the allowed margin below the current version's.
On the worked example: 0.90 < 0.95 − 0.02 = 0.93, so roll back.
Level 3: in Python
p_canary, p_control, delta = 0.90, 0.95, 0.02
round(p_control - delta, 2) # → 0.93
# roll back?
p_canary < p_control - delta # → True
Users are assigned to the canary by hashing their id into one of 100 buckets: the same user always lands in the same bucket, so nobody flips between versions mid-conversation, and growing from 5% to 25% only adds users.
Figure 4 · Diagram
sequenceDiagram participant D as Deploy system participant R as Router participant M as Metrics D->>R: send 1% of users to v2 R->>M: success rates, v1 vs v2 M-->>D: v2 95%, v1 95%: healthy D->>R: grow to 5% R->>M: success rates, v1 vs v2 M-->>D: v2 91%, v1 95%: worse by 4 points D->>R: roll back: 0% to v2 D-->>D: alert the owner with traces
Figure 5 · Chart
The good version tracks the control through 1, 5, 25, 50 and 100%; the bad one passes the 1% step by luck and is rolled back at 5%
In code: in_canary hashes a user id into one of 100 buckets.
CanaryController.observe accumulates success counts, and
CanaryController.step applies the rollback rule above: advance, wait for
more samples, or roll back. simulate_rollout drives a whole rollout for the
figure.
A feature flag is the same idea as an on/off switch in configuration: it turns a capability on for chosen tenants or users without redeploying, and off again just as fast.
Chapter 4
Kill switches and rate limits: bound the blast radius
Everyday picture The red emergency-stop button on a treadmill, and the turnstile at a stadium that lets people through only so fast however hard the crowd pushes.
A kill switch stops the agent instantly, for one tenant (one customer organisation) or for everyone. A rate limit caps actions per unit time, so even a malfunctioning agent can only send so many emails or change so many records per minute. The blast radius is how much damage a failure can do before someone notices; these two mechanisms bound it.
Worked example: the token bucket. Picture a jar that holds at most 5 tokens and gains 1 token per second. Each action takes a token; with no token, the action is refused. Starting full: 5 actions in a burst succeed, the 6th is refused. After 2 seconds the jar has 2 tokens again: 2 more succeed, the next is refused.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| tokens in the bucket at time | 2 at t = 2 s | |
| when the bucket was last updated | 0 s | |
| tokens left then | 0 after the burst | |
| refill rate, tokens per second | 1 | |
| capacity: the largest burst allowed | 5 | |
| the smaller of the two: the jar can't overflow |
In words: the tokens now are what was left plus what has dripped in since, but never more than the jar holds.
On the worked example: min(5, 0 + 1 × (2 − 0)) = 2 tokens, so two more actions are allowed.
Level 3: in Python
# capacity, tokens per second
C, r = 5, 1
def b(t, t_0, b_t0):
# what was left plus the drip, capped at C
return min(C, b_t0 + r * (t - t_0))
# 2 s after the burst emptied the jar
b(2, 0, 0) # → 2
# a long wait refills only to capacity
b(60, 0, 0) # → 5
Figure 6 · Chart
A burst of 8 requests gets 5 through before the bucket empties, then requests pass one per second at the refill rate
In code: KillSwitch holds the global and per-tenant off switches.
TokenBucket.level is the formula for , and TokenBucket.allow
spends a token or refuses.
Chapter 5
A tamper-evident audit log
Everyday picture A ledger where every page ends with a wax seal pressed from the previous page's seal plus this page's contents. Change one number on page 12 and its seal no longer matches, and neither does any seal after it.
An audit log records which agent took which action, for which user, with what inputs. For regulated work, it must be tamper-evident: an alteration can't go unnoticed. A hash chain does this: each entry stores the hash of the previous entry, and its own hash covers its contents and that previous hash.
Level 3: the formula and its symbols
(Entries are numbered from 1 here; verify() reports Python list
positions, which start at 0, so "entry 2" is position 1.)
Symbols
| Symbol | Meaning here |
|---|---|
| the fingerprint (hash) stored with entry | |
| the previous entry's fingerprint | |
| the hash function: any input to a 64-hex-digit fingerprint | |
| "joined together with" (concatenation) | |
| a fixed starting value for the first entry |
In words: each entry's fingerprint is the fingerprint of the previous fingerprint joined to this entry's contents.
On the worked example: three entries: refund A200 for 40, refund A201
for 15, escalate A300. Change entry 2's amount from 15 to 1,500 and
recomputing its fingerprint no longer gives the stored , so verify()
reports position 1. Delete entry 2 instead, and entry 3 moves up to position
1; its stored previous fingerprint is , the deleted entry's, which
doesn't match , so the break is again reported at position 1.
In Python:
import hashlib
def seal(prev, entry):
# SHA256(h_{i-1} ‖ entry_i)
return hashlib.sha256((prev + entry).encode()).hexdigest()
# h_0
log, prev = [], "0" * 64
for entry in ["refund A200 40", "refund A201 15", "escalate A300"]:
prev = seal(prev, entry)
# each entry is stored with its h_i
log.append((entry, prev))
def verify(log):
prev = "0" * 64
for position, (entry, h_i) in enumerate(log):
if seal(prev, entry) != h_i:
# the first broken link
return position
prev = h_i
verify(log) is None # → True
# entry 2's amount changed
verify([log[0], ("refund A201 1500", log[1][1]), log[2]]) # → 1
# entry 2 deleted
verify([log[0], log[2]]) # → 1
Figure 7 · Diagram
flowchart LR G["h0 = 000…"] --> E1["entry 1: refund A200, 40<br/>prev = h0<br/>h1 = SHA256(h0 ‖ entry 1)"] E1 --> E2["entry 2: refund A201, 15<br/>prev = h1<br/>h2 = SHA256(h1 ‖ entry 2)"] E2 --> E3["entry 3: escalate A300<br/>prev = h2<br/>h3 = SHA256(h2 ‖ entry 3)"]
In code: AuditLog.append stores the previous entry's hash in the new
entry and seals it with its own; AuditLog.verify walks the chain and
returns the position of the first broken link, or None.
Chapter 6
Change management: adoption is part of the job
A technically working agent still fails if people don't trust or use it. Involve the people whose work it touches from shadow mode onwards (their corrections are your best eval data), design around their actual workflow, show sources and confidence so they can check answers, and make handing a case to a person one click.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1How would you roll out an agent that takes real actions in a company's ERP system (its finance and operations system of record)?Think it through, then reveal
Start read-only in shadow mode on real transactions, comparing proposals with what staff actually did, broken down by action type. Give the agent a dedicated service identity with the narrowest permissions, a sandbox or dry-run mode for writes, and idempotency keys so retries can't double-post. Move to proposing actions that a person approves in their existing workflow, starting with the low-value, reversible action types that scored best in shadow. Grant autonomy only for those, with value thresholds that still require approval, per-tenant kill switches, rate limits on writes, a canary rollout for every prompt or model change, eval gates in CI, and a tamper-evident audit log tied to traces. Demote automatically on incidents or drift, and keep the finance team involved throughout.
Question 2Why start in shadow mode instead of with human approval?Think it through, then reveal
Shadow mode costs the humans nothing extra and produces a clean comparison against what they actually did, at full volume, before the agent can influence anything. Approval mode adds review load and can bias reviewers toward accepting the agent's suggestion.
Question 3What does a canary protect against that offline evals don't?Think it through, then reveal
Real traffic: inputs, integrations and user behaviour the golden set didn't anticipate. Evals gate the release; the canary limits the damage from whatever the evals missed.
Question 4Why a hash chain rather than an ordinary log table?Think it through, then reveal
Anyone with write access can quietly edit an ordinary row. With a hash chain, any edit or deletion breaks every later link, so tampering is detectable, and anchoring the latest hash somewhere separate (or using write-once storage) makes it provable.
Primary sources
The papers behind this lesson
Haber & Stornetta, How to Time-Stamp a Digital Document (Journal of Cryptology, 1991), Introduced chaining each record to the hash of the one before, the idea behind tamper-evident logs (and, later, blockchains).
The paper ↗Researcher's shelf
Further reading
- Google SRE workbook, canarying releases: https://sre.google/workbook/canarying-releases/
- Martin Fowler, feature toggles: https://martinfowler.com/articles/feature-toggles.html
- Token bucket algorithm: https://en.wikipedia.org/wiki/Token_bucket
- NIST AI Risk Management Framework: https://www.nist.gov/itl/ai-risk-management-framework
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.