At a glance
Key takeaways
- Tokenize the prompt, look up a vector per token, add position, run the transformer blocks, score every vocabulary entry from the last position, softmax, sample, append, repeat.
- Temperature divides the scores before softmax: low is predictable, high is varied.
- The naive loop re-reads everything each step; the KV cache makes each step cost one token. Training is the same forward pass plus a next-token loss.
Level 2
How it works, from scratch
This lesson wires the real pieces from the other lessons into one working
pipeline: the tokenizer from primer.ml.tokenization, the transformer from
primer.ml.transformer, and a sampling loop. Keep this one picture in your
head; every other lesson zooms into one box of it.
Figure 1 · Diagram
flowchart LR A[Prompt text] --> B[Tokenizer<br/>text to IDs] B --> C[Embedding lookup<br/>IDs to vectors] C --> D[Add position info] D --> E[Transformer blocks<br/>repeated N times] E --> F[Output layer<br/>score per vocab token] F --> G[Softmax + sampling] G --> H[Next token] H -->|append and repeat| E
Chapter 1
Text to ids: the coat check
Everyday picture A coat check. You hand over a coat (a piece of text) and get back a numbered ticket (a token id). The model only ever handles the tickets.
Tiny worked example With this module's toy tokenizer, "Reset your
password" becomes 5 tickets: Re set y our password →
[346, 377, 309, 328, 291]. The common word " password" got a single
ticket; the rarer pieces got several.
Figure 2 · Diagram
flowchart LR T["'Reset your password'"] --> TK["tokenizer<br/>(learned kit of 400 pieces)"] TK --> I["[346, 377, 309, 328, 291]"]
The code tok.encode(prompt); how the kit is learned is the whole of
primer.ml.tokenization.
In code: build_pipeline trains the toy tokenizer (primer.ml.tokenization.ByteBPE) and builds an untrained primer.ml.transformer.TinyGPT sized to its vocabulary.
Why it matters Prompt length, price and context limits are all counted in these tickets.
Chapter 2
Ids to vectors: a lookup, not a computation
Everyday picture A dictionary where ticket number 291 opens to page 291, and each page holds a list of numbers describing that token. A second dictionary, indexed by seat number, describes where the token sits.
Tiny worked example The table has 400 rows (one per token) of 32 numbers. Id 291 fetches row 291. The 5 ids fetch 5 rows: a 5 × 32 grid. Row i of the position table (i = 0 to 4) is added to row i of that grid.
Figure 3 · Diagram
flowchart LR
I["ids (5)"] --> E["token table<br/>400 × 32"]
E --> X["5 × 32: what each token is"]
P["positions 0..4"] --> PT["position table<br/>64 × 32"]
PT --> Y["5 × 32: where each token is"]
X --> ADD(("+"))
Y --> ADD
ADD --> OUT["5 × 32 input to the blocks"]
primer.ml.embeddings).The math and the code
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| a token's position in the prompt, from 0 | 4 (the last token) | |
| the token id at position | = 291 | |
| the token embedding table | 400 × 32 | |
| row of that table | row 291: 32 numbers | |
| row of the position table | row 4: 32 numbers | |
| what the first block receives for position | 32 numbers |
In words: "each token's input vector is its token's row plus its position's row."
With the numbers: , one 32-number list plus
another. The code is model.wte[ids] + model.wpe[:len(ids)]. In miniature,
with 3-number rows instead of 32: = (0.2, −0.1, 0.5) and =
(0.1, 0.3, −0.2) give = (0.3, 0.2, 0.3).
Level 3: in Python
# row 291 of the token table (3 numbers, not 32)
E_291 = [0.2, -0.1, 0.5]
# row 4 of the position table
P_4 = [0.1, 0.3, -0.2]
# E_(t_i) + P_i, number by number
x_4 = [round(e + p, 2) for e, p in zip(E_291, P_4)]
x_4 # → [0.3, 0.2, 0.3]
In code: trace does both lookups and the add, and keeps every stage's result so you can inspect the grid before and after positions are mixed in.
Why it matters This is the only place a token's identity enters the model; every later step works on these vectors.
Chapter 3
The transformer blocks: rounds of meeting and desk work
Everyday picture The team from primer.ml.transformer: each round is a
meeting where every token listens to the others (attention), then desk work
where each token thinks alone (feed-forward). This toy runs 2 rounds; large
models run dozens.
Tiny worked example The 5 × 32 grid goes into block 1 and comes out 5 × 32; the same through block 2. By the end, the vector at the last position ("password") has absorbed information from "Re", "set", "y" and "our".
Figure 4 · Diagram
flowchart LR X["5 × 32"] --> B1["block 1<br/>meeting + desk work"] --> B2["block 2"] --> LN["final norm"] --> H["5 × 32<br/>context-aware vectors"]
The math and the code
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the 5 × 32 input grid from section 2 | ||
| the -th transformer block | = 2 | |
| "and so on, for every block in between" | ||
| the final layer norm | ||
| the context-aware vectors | 5 × 32 |
In words: "run the input through every block in turn, then normalize."
With the numbers: here N = 2, so h = LN(Block₂(Block₁(x))). The code is
the for block in model.blocks loop in trace. In miniature, with one
3-number vector and two stand-in blocks that each add an edit: x = (1, 2, 3),
Block₁ adds (0, 1, 0) and Block₂ adds (0, 0, 2), giving (1, 3, 5). The final
norm subtracts the mean (3) and divides by the spread (1.63), so
h = (−1.22, 0, 1.22).
Level 3: in Python
import statistics
x = [1.0, 2.0, 3.0]
# a stand-in edit
def block_1(v): return [v_i + d for v_i, d in zip(v, [0, 1, 0])]
def block_2(v): return [v_i + d for v_i, d in zip(v, [0, 0, 2])]
def LN(v):
mu, sigma = statistics.fmean(v), statistics.pstdev(v)
# centre, then rescale
return [round((v_i - mu) / sigma, 2) for v_i in v]
h = x
# Block_1 first, then Block_2 ... up to Block_N
for block in [block_1, block_2]:
h = block(h)
h # → [1.0, 3.0, 5.0]
LN(h) # → [-1.22, 0.0, 1.22]
Why it matters This is where nearly all the compute and all the "understanding" happen.
Chapter 4
Scores, softmax and temperature
Everyday picture A scoreboard with one line for every token in the vocabulary. Softmax turns the scores into shares of a pie. Temperature sets how adventurous the pick is: low temperature almost always takes the biggest slice; high temperature gives the small slices a real chance.
Tiny worked example Three candidate tokens scored 2.0, 1.0 and 0.5.
| temperature | p(A) | p(B) | p(C) |
|---|---|---|---|
| 0 (greedy) | 1.000 | 0.000 | 0.000 |
| 0.5 | 0.844 | 0.114 | 0.042 |
| 1 | 0.629 | 0.231 | 0.140 |
| 2 | 0.481 | 0.292 | 0.227 |
Figure 5 · Diagram
flowchart LR H["last row of h<br/>32 numbers"] --> S["× token tableᵀ<br/>400 scores"] S --> T["÷ temperature"] T --> SM["softmax<br/>400 probabilities"] SM --> PICK["draw one token"]
The math and the code Softmax raises e (≈ 2.718) to the power of each score, so bigger scores get disproportionately bigger shares, then divides by the total so the shares add up to 1:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the score of candidate token | = 2.0 | |
| the temperature | 0.5 | |
| e ≈ 2.718 raised to that power | = 54.6 | |
| number of candidates (the vocabulary size) | 3 here, 400 in the model | |
| add up over every candidate | ||
| probability that token comes next | 0.844 |
In words: "the chance of token i is e to the power of its score over the temperature, divided by the same quantity summed over all tokens."
With the numbers: T = 0.5 doubles every score to (4, 2, 1): e⁴ = 54.6,
e² = 7.39, e¹ = 2.72, total 64.7, so p(A) = 54.6 / 64.7 = 0.844
(next_token_probs). Temperature 0 is the limit case: all probability on the
top score.
Level 3: in Python
import math
z = [2.0, 1.0, 0.5]
T = 0.5
# e^(z_i / T) for each candidate
exps = [math.exp(z_i / T) for z_i in z]
[round(e, 2) for e in exps] # → [54.6, 7.39, 2.72]
# Σ_j e^(z_j / T)
total = sum(exps)
round(total, 1) # → 64.7
# p_i: each share of the total
[round(e / total, 3) for e in exps] # → [0.844, 0.114, 0.042]
Figure 6 · Drawn from the lesson's code
At T = 0.25 token A takes 98% of the probability; at T = 2 the three tokens share it 48%, 29% and 23%
Figure 7 · Drawn from the lesson's code
The untrained model's 12 favourites are random byte fragments, each only 1.25 to 1.4 times the uniform 1-in-400 share
Why it matters Use low temperature for extraction and tool calls, where you want the most likely answer, and higher temperature for creative writing. Temperature 0 reduces randomness but doesn't guarantee identical outputs on real serving hardware.
Chapter 5
The loop: append and repeat
Everyday picture Writing a sentence one word at a time, and re-reading everything you've written before choosing each new word.
Tiny worked example A 2-token prompt, 4 new tokens. Step 1 reads 2 tokens, step 2 reads 3, step 3 reads 4, step 4 reads 5: 14 token positions to produce 4 tokens. With a 20-token prompt and 200 new tokens, the naive loop reads 23,900 positions; a KV cache reads 219.
Figure 8 · Diagram
flowchart LR S["ids so far"] --> M["full forward pass<br/>over ALL ids"] M --> P["probabilities for the next id"] P --> D["draw one id"] D --> A["append it"] A -->|"not done"| S A -->|"stop token or length limit"| OUT["decode ids to text"]
primer.ml.inference) stores it once.The math Total positions processed without a cache:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| prompt length in tokens | 2 | |
| tokens to generate | 4 | |
| the step counter, from 0 to | 0, 1, 2, 3 | |
| tokens re-read at step | 2, 3, 4, 5 | |
| total positions run through the model | 14 |
In words: "each step re-reads the prompt plus everything generated so far; add that up over all steps."
With the numbers: 4 × 2 + (4 × 3) / 2 = 8 + 6 = 14, the count
generate reports as positions_processed.
Level 3: in Python
p, n = 2, 4
# Σ over t = 0 .. n-1 of (p + t): 2 + 3 + 4 + 5
sum(p + t for t in range(n)) # → 14
# the closed form gives the same count
n * p + n * (n - 1) // 2 # → 14
Figure 9 · Drawn from the lesson's code
Without a cache work grows with the square of output length, 23,900 positions at 200 tokens against 219 with a KV cache
In code: Generation holds the generated ids, the decoded text and the positions count; naive_vs_cached_work counts the positions processed with and without a KV cache for the figure.
Why it matters Output length drives latency and cost; this is why every serving system caches keys and values.
Chapter 6
Training: the same forward pass, plus a loss
Everyday picture A guessing game with instant feedback. Cover the next word, guess it, uncover it, and note how surprised you were. Training nudges every weight to make the surprise smaller next time.
Tiny worked example If the model gives the right next token probability 0.5, the loss is −ln 0.5 = 0.693. Our untrained model averages 5.97 on a real sentence, close to ln 400 = 5.99, the score of a blind uniform guess.
Figure 10 · Diagram
flowchart LR T["training text"] --> F["forward pass<br/>(this whole lesson)"] F --> P["probabilities at every position"] P --> L["loss: −ln p(actual next token)<br/>averaged over positions"] L --> B["backpropagation<br/>gradient for every weight"] B --> U["optimizer nudges weights"] U -->|next batch| F
primer.ml.neural_net) works out how each
weight contributed, and the optimizer (primer.ml.optimizers) adjusts them.
One pass over n tokens gives n − 1 guesses at once, in parallel, which is a
big reason transformers train fast.The math and the code A logarithm answers "to what power must I raise e to get this number?" For probabilities between 0 and 1 it is negative, so we flip the sign; ln 1 = 0 (no surprise) and ln of a tiny number is very negative (huge surprise).
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| tokens in the training text | 12 ("The model reads tokens, not words.") | |
| the token at position | ||
| probability the model gave the actual next token, having seen everything before it; the bar reads "given" | ≈ 1/400 when untrained | |
| natural logarithm | ln(1/400) = −5.99 | |
| the average over all guesses | ||
| the loss training pushes down | 5.97 |
In words: "for every position, take the log of the probability the model gave the true next token, average them, and flip the sign."
With the numbers: a uniform guess over 400 tokens gives every true token
p = 1/400, so each term is −ln(1/400) = ln 400 = 5.99. Our untrained model
scores 5.97, so it is still essentially guessing. Perplexity, e raised to
the loss, is 393: "as unsure as choosing among 393 equally likely tokens"
(next_token_loss, cross_entropy).
Level 3: in Python
import math
# one guess that gave the true token p = 0.5
round(-math.log(0.5), 3) # → 0.693
n = 12
# a uniform guess gives every true token 1/400
p = [1 / 400] * (n - 1)
# -(1/(n-1)) Σ ln p
L = -sum(math.log(p_i) for p_i in p) / (n - 1)
round(L, 2) # → 5.99
Why it matters Pretraining is exactly this, over trillions of tokens: predicting the next token well forces the model to absorb grammar, facts and reasoning patterns. Random weights produce gibberish, and training is what turns the same machinery into a useful model.
Test yourself
6 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1What happens, step by step, when you send a prompt?Think it through, then reveal
Tokenizer turns text into ids; each id looks up an embedding; position information is added; the vectors pass through N transformer blocks; the last position's vector is scored against the whole vocabulary; softmax and sampling pick a token; it's appended and the loop repeats until a stop token.
Question 2Why does only the last position matter when generating?Think it through, then reveal
Its vector has attended to every earlier token and is the one trained to predict what comes next; earlier rows predict tokens we already have.
Question 3What does temperature do to the scores (2, 1, 0.5) at T = 0.5?Think it through, then reveal
Doubles them to (4, 2, 1) before softmax, sharpening the distribution: the top token rises from 0.63 to 0.84.
Question 4How many token positions does the naive loop process for a 2-token prompt and 4 new tokens?Think it through, then reveal
2 + 3 + 4 + 5 = 14. A KV cache avoids re-reading the unchanged prefix.
Question 5What loss does an untrained model get over a 400-token vocabulary, and why?Think it through, then reveal
About ln 400 ≈ 5.99 (perplexity about 400), because with random, tiny weights its predictions are close to uniform.
Question 6How is training different from inference?Think it through, then reveal
Same forward pass, plus a loss comparing predictions to the real next tokens, then backpropagation and a weight update. Inference only runs the forward pass.
Primary sources
The papers behind this lesson
The architecture inside the "transformer blocks" box.
Read on rumblr →The paper ↗The decoder-only, next-token-prediction recipe this pipeline follows, including learned positions, byte-level BPE and tied embeddings.
The paper ↗Showed that scaling this same loop up produces models that follow instructions from examples in the prompt.
Read the annotated companion →The paper ↗Introduced nucleus (top-p) sampling and explained why pure greedy decoding produces repetitive text.
The paper ↗Researcher's shelf
Further reading
- Andrej Karpathy, Let's build GPT (video): https://www.youtube.com/watch?v=kCc8FmEb1nY
- Karpathy's
nanoGPT: https://github.com/karpathy/nanoGPT - Jay Alammar, The Illustrated GPT-2: https://jalammar.github.io/illustrated-gpt2/
- 3Blue1Brown, But what is a GPT? (video): https://www.youtube.com/watch?v=wjZofJX0v4M
- Hugging Face, How to generate text (decoding strategies): https://huggingface.co/blog/how-to-generate
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.