At a glance
Key takeaways
- Measure cost per successful task, not per call; include retries and human cleanup.
- Free wins first: prompt caching (stable prefix first), trimming tool output and chunks, concise output, batch APIs for non-interactive work.
- Then route easy tasks to small models, validated with evals.
- Run independent tool calls in parallel; stream to improve perceived speed.
- Budgets per task, spend caps per tenant, and alerts on sudden jumps in tokens per task.
- Semantic caches can serve confidently wrong answers; use them only on narrow traffic, with guards and a measured wrong-hit rate.
Level 2
How it works, from scratch
A restaurant doesn't cut costs by buying worse ingredients for every dish. It sends simple orders to the line cook and complex ones to the head chef, pre-chops what every dish shares, doesn't plate food nobody eats, cooks things in parallel, does bulk prep overnight when it's cheaper, and watches the one kitchen station whose bills suddenly spike.
Every lever in this lesson is one of those moves. The measure that matters is cost per successful task, not cost per call: a cheap attempt that fails still has to be paid for, and so does fixing it.
All prices in this module are illustrative constants chosen for round arithmetic (a "large" model at $5 per million input tokens and $25 per million output tokens; a "small" one at $1 and $5). They show the shape of the trade-offs; check your provider's current price list for real numbers.
Chapter 1
How a request is priced
Everyday picture A taxi that charges one rate for the distance to your
pickup and a higher rate for the ride itself. Models charge per token
(a word piece, about 4 characters of English): one rate for tokens you send
(input) and a higher rate for tokens the model writes (output),
because writing happens one token at a time (see primer.ml.inference).
Worked example A large-model call with 10,000 input tokens and 500 output tokens: 10,000 × $5 / 1,000,000 = $0.05 for input, plus 500 × $25 / 1,000,000 = $0.0125 for output, total $0.0625. If 8,000 of those input tokens are a stable prefix served from the prompt cache (reused work from an earlier identical beginning, billed here at 10% of the input price), input becomes 8,000 × $0.50/M + 2,000 × $5/M = $0.014, and the call costs $0.0265, 58% less.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| input tokens sent | 10,000 | |
| input tokens served from the prompt cache | 8,000 | |
| output tokens generated | 500 | |
| price per million input / output tokens | $5, $25 | |
| cached-read price as a fraction of normal input price (Greek rho) | 0.1 | |
| prices are per million tokens |
In words: cached input at its discounted rate, plus the rest of the input at full rate, plus output at the output rate, all per million.
On the worked example: (8,000 × 0.1 × 5 + 2,000 × 5 + 500 × 25) / 10⁶ = (4,000 + 10,000 + 12,500) / 10⁶ = $0.0265.
Level 3: in Python
n_in, c, n_out = 10_000, 8_000, 500
p_in, p_out, rho = 5, 25, 0.1
cost = (c * rho * p_in + (n_in - c) * p_in + n_out * p_out) / 10**6
round(cost, 4) # → 0.0265
Two facts fall out: output tokens cost several times more than input (ask
for concise answers), and a cached prefix is nearly free (put stable content
first; see primer.agents.context). Caches usually charge a small premium
the first time a prefix is written; this module ignores it for simplicity.
In code: request_cost is the formula above, reading each model's
Price from PRICES.
Chapter 2
Route each task to the cheapest model that can do it
Everyday picture A hospital triage nurse: sprained ankles go to the nurse practitioner, chest pains to the cardiologist. Nobody sends every patient to the most expensive specialist.
Worked example The sample workload is 10 tasks: 7 simple (classify a ticket, extract a date) and 3 complex (plan a migration). All on the large model it costs $0.671. Routing the 7 simple ones to the small model, which is 5x cheaper per token, brings it to $0.314: 53% saved, with the hard tasks still on the strong model.
Figure 1 · Diagram
flowchart LR
T[Incoming task] --> R{Router<br/>rules or a small classifier}
R -->|classify, extract,<br/>format, short| S[Small, fast model]
R -->|plan, analyze,<br/>multi-step reasoning| L[Large model]
S --> O[Result]
L --> O
primer.agents.evals): a router that sends hard tasks to the small model
saves money and quietly loses quality.In code: route is the rule-based router; workload_cost prices a list
of WorkItem tasks with any levers switched on; routing_savings compares
all-large against routed on SAMPLE_WORKLOAD.
Chapter 3
Response caching and semantic caching
Everyday picture A receptionist who's been asked "what's the wifi
password?" a hundred times just answers from memory. That's an exact
response cache: the same question (ignoring case and spaces) returns the
stored answer with no model call. A semantic cache goes further and
answers similar questions from memory, using embeddings (vectors where
closeness means similar meaning, see primer.ml.embeddings) to decide
"similar". That's where the receptionist starts giving the vacation-policy
answer to someone asking about sick leave.
Worked example With the toy embedder in primer.common.embedder:
| Stored question | New question | Similarity | Same intent? |
|---|---|---|---|
| How do I reset my password? | I forgot my password, how do I recover it? | 0.88 | yes |
| How many vacation days do I get? | How many sick days do I get? | 0.97 | no |
| Request a new laptop | Request a new monitor | 0.96 | no |
| What does ERR-4012 mean? | What does ERR-4013 mean? | 0.56 | no |
The two near-misses score higher than a genuine paraphrase. No single threshold separates them. Two cheap guards help: identifiers and numbers must match exactly (ERR-4012 ≠ ERR-4013), and negation must match ("cancel" ≠ "don't cancel"). They can't catch "sick" vs "vacation", which share a topic and differ in a single word.
Figure 2 · Diagram
flowchart TD
Q[New question] --> E{Exact match<br/>in response cache?}
E -->|yes| A1[Return stored answer]
E -->|no| V[Embed and find the<br/>nearest stored question]
V --> T{Similarity above<br/>threshold?}
T -->|no| M[Call the model]
T -->|yes| G{IDs, numbers and<br/>negation identical?<br/>entry not expired?}
G -->|no| M
G -->|yes| A2[Return stored answer]
M --> S[Store the new answer]
Figure 3 · Chart
No threshold separates them: wrong hits stay at 40 to 60% through 0.95, often above the paraphrase hit rate, and reach 0 only where paraphrase hits do too
In code: ResponseCache is the exact cache. SemanticCache is the
semantic path: SemanticCache.lookup skips expired entries and entries whose
identifiers or negation differ, then returns the nearest survivor, and
SemanticCache.get applies the threshold. semantic_cache_sweep draws the
figure from the labelled pairs in CACHE_PAIRS.
Chapter 4
Trim tokens
Everyday picture Don't photocopy the whole binder when the colleague
needs one page. Compress tool results to the fields the next step needs
(primer.agents.context.compress_tool_output), send the top few reranked
chunks instead of dozens, remove repeated boilerplate from system prompts,
and ask for concise output, because output is the expensive direction.
In code: workload_cost models trimming as cutting each task's tool
output to 500 tokens, and concise output as cutting long answers by a third.
Chapter 5
Run independent tool calls at the same time
Everyday picture Boil the pasta while the sauce simmers. Cooking them one after the other takes the sum of the times; together, the time of the slowest.
Worked example Three independent lookups of 100 ms each: sequentially about 300 ms; concurrently about 100 ms.
Figure 4 · Diagram
sequenceDiagram
participant A as Agent
participant T1 as Weather API
participant T2 as Calendar API
participant T3 as CRM API
Note over A,T3: Sequential: about 300 ms
A->>T1: call
T1-->>A: result
A->>T2: call
T2-->>A: result
A->>T3: call
T3-->>A: result
Note over A,T3: Parallel: about 100 ms
par
A->>T1: call
and
A->>T2: call
and
A->>T3: call
end
T1-->>A: result
T2-->>A: result
T3-->>A: result
Figure 5 · Chart
Three 100 ms tool calls take 300 ms end to end when run one after another, but only 100 ms when started together
asyncio: the real timings land within a few milliseconds
of these bars.In code: run_sequential awaits each call before starting the next;
run_parallel starts them all with Python's asyncio gather and waits once.
Streaming (showing tokens as they're generated) doesn't reduce total time either, but users see progress immediately, which changes how fast the system feels.
Chapter 6
Batch what isn't interactive
Everyday picture Sending the laundry out to be done overnight at half price instead of waiting at the express counter. Providers offer batch APIs: submit many requests, get results within hours (often within 24), typically at about half the price. Use them for anything no one is waiting on: nightly evals, backfilling document processing, bulk classification.
In code: batch_cost prices a list of requests at the batch discount.
Chapter 7
Budgets and alerts
Everyday picture A prepaid card with a hard limit, and a bank that texts you when a purchase looks nothing like your usual spending.
A task budget caps steps and tokens per task, so a confused agent stops instead of looping all night. A tenant spend cap (a tenant is one customer organisation on a shared platform) stops one customer's runaway usage from becoming a surprise invoice. An anomaly alert fires when a task uses far more tokens than usual:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| tokens used by the task just finished | 10,000 | |
| mean tokens per task over this tenant's recent history (Greek mu) | 1,000 | |
| standard deviation of that history: the typical distance from the mean (Greek sigma) | ≈ 71 | |
| how many standard deviations count as unusual | 3 |
In words: alert when a task uses more tokens than the usual amount plus three times the usual spread.
On the worked example: history 1000, 1100, 900, 1050, 950, 1000 has mean 1,000 and standard deviation ≈ 71 (the squared distances from the mean add up to 25,000; divided by 6 − 1 = 5 that's 5,000, whose square root is 70.7), so the line is 1,000 + 3 × 70.7 ≈ 1,212. A 10,000-token task is far above it: alert. Such jumps usually mean a loop or a bad deploy.
Level 3: in Python
import statistics
history = [1000, 1100, 900, 1050, 950, 1000]
mu = statistics.mean(history)
# divides by 6 - 1, as above
sigma = statistics.stdev(history)
mu, round(sigma, 1) # → (1000, 70.7)
z = 3
# the alert line
round(mu + z * sigma) # → 1212
# alert?
10_000 > mu + z * sigma # → True
In code: TaskBudget.charge counts each step's tokens and raises
BudgetExceeded at either limit; TenantSpend.record adds a finished task's
cost to its tenant's total and returns a cap alert or an anomaly alert (the
formula above).
Chapter 8
Unit economics: cost per successful task
Everyday picture A cheap printer that jams on 40% of pages isn't cheap once you count the wasted paper and the time spent clearing jams.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Small model | Large model |
|---|---|---|---|
| model cost of one attempt | $0.002 | $0.010 | |
| chance an attempt succeeds | 0.60 | 0.95 | |
| expected number of attempts until one succeeds | 1.67 | 1.05 | |
| cost of a person fixing one failure (illustrative: a few minutes of staff time) | $2.00 | $2.00 |
In words: if you can simply retry, you pay one attempt's cost for every expected attempt; if failures need a person, every task pays its attempt plus the chance of failure times the cost of the fix.
On the worked example: retrying, the small model costs 0.002 / 0.60 = $0.0033 per success and the large one 0.010 / 0.95 = $0.0105, so the small model wins. With human cleanup, the small model costs 0.002 + 0.40 × 2.00 = $0.802 per task and the large one 0.010 + 0.05 × 2.00 = $0.110: the "cheap" model is 7x more expensive.
Level 3: in Python
def per_success(c, p):
# retry until it works
return c / p
def per_task(c, p, h=2.00):
# a person fixes each failure
return c + (1 - p) * h
round(per_success(0.002, 0.60), 4), round(per_success(0.010, 0.95), 4) # → (0.0033, 0.0105)
round(per_task(0.002, 0.60), 3), round(per_task(0.010, 0.95), 3) # → (0.802, 0.11)
# small vs. large, with cleanup
round(per_task(0.002, 0.60) / per_task(0.010, 0.95)) # → 7
Figure 6 · Chart
With automatic retries the small model is cheaper per success (0.3 vs 1.1 cents); if a person fixes each failure the large model wins, 11 cents vs 80
In code: cost_per_success_with_retries is and
cost_per_success_with_cleanup is .
Chapter 9
Putting it together: a 5x plan, in order
Order the levers so the ones that can't hurt quality come first:
- Prompt caching: reorder the prompt so the stable prefix is reused. No behaviour change.
- Trim tokens: compress tool results and send fewer chunks.
- Concise output: ask for shorter answers where length adds nothing.
- Route easy tasks to the small model: the first lever that can change quality, so it's validated with evals before and after.
- Batch the non-interactive share.
Figure 7 · Chart
Workload cost falls from 67 to 8 cents as levers stack: caching 2.2x, trimming 3.6x, concise output 3.9x, routing 7.8x, batching 8.4x
In code: five_x_plan switches the levers on one at a time, in this
order, and reports the cost and cumulative reduction after each.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1How would you cut the cost per task by 5x without hurting quality?Think it through, then reveal
First measure: cost per successful task, broken down by step, model, input vs. output, and cached vs. uncached tokens, with an eval set to hold quality fixed. Then, in order: restructure prompts for prompt caching; compress tool results and send only the top reranked chunks; ask for concise output; route simple steps (classification, extraction, formatting) to a small model, checking the eval before and after; move anything non-interactive to a batch API; cap steps and tokens per task. On the sample workload here that's about 8x, and every step after the first three is checked against the eval set.
Question 2Why can a cheaper model cost more?Think it through, then reveal
Because failures cost money too: retries, human cleanup, lost customers. At 60% success with a $2 human fix, a $0.002 call costs $0.80 per task; a $0.010 call at 95% costs $0.11.
Question 3What are the risks of a semantic cache?Think it through, then reveal
Returning a confident, wrong answer to a question that's similar but not the same ("sick days" vs "vacation days"), and serving stale answers after facts change. Mitigate with high thresholds, exact-match guards on IDs, numbers and negation, per-intent scoping, a TTL, and a measured wrong-hit rate on labelled pairs before enabling it.
Question 4Your average tokens per task doubled overnight. What do you check?Think it through, then reveal
Whether a deploy changed a prompt or a tool (a bigger tool output, a new
retrieval setting), whether an agent is looping (the same tool called with
the same arguments), and whether the prompt cache hit rate dropped
(something volatile moved to the top). Traces make this a lookup rather
than a guess (primer.agents.observability).
Primary sources
The papers behind this lesson
Hinton, Vinyals & Dean, Distilling the Knowledge in a Neural Network (2015), Showed how to train a small model to imitate a large one, the technique behind making the cheap model good enough to take more of the routed traffic.
Read the annotated companion →The paper ↗Chen, Zaharia & Zou, FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance (2023), Studied prompt adaptation, caching and cascades that try cheap models first and escalate to expensive ones only when needed.
The paper ↗Researcher's shelf
Further reading
- Anthropic prompt caching docs: https://docs.claude.com/en/docs/build-with-claude/prompt-caching
- Anthropic Message Batches docs: https://docs.claude.com/en/docs/build-with-claude/batch-processing
- Anthropic, parallel tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
- Python
asyncio.gather: https://docs.python.org/3/library/asyncio-task.html#asyncio.gather
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.