rumblr Work in progressWIP

● The AI Primer · Lesson 1 · Part 1: how the model works inside

The big picture

what happens when you send a prompt

This lesson covers What happens, end to end, when you send a prompt

Free lesson 24 min10 figures and diagrams6 interactive
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Tokenize the prompt, look up a vector per token, add position, run the transformer blocks, score every vocabulary entry from the last position, softmax, sample, append, repeat.
  2. Temperature divides the scores before softmax: low is predictable, high is varied.
  3. The naive loop re-reads everything each step; the KV cache makes each step cost one token. Training is the same forward pass plus a next-token loss.

Level 2

How it works, from scratch

This lesson wires the real pieces from the other lessons into one working pipeline: the tokenizer from primer.ml.tokenization, the transformer from primer.ml.transformer, and a sampling loop. Keep this one picture in your head; every other lesson zooms into one box of it.

Figure 1 · Diagram

Reading it: read left to right, then follow the loop back. The prompt is cut into token ids, each id picks a vector from a table, and position information is mixed in so order counts. The vectors pass through the transformer blocks, where tokens exchange information. The output layer turns the last token's final vector into a score for every token in the vocabulary; softmax makes those scores probabilities and one token is drawn. That token is appended and the loop runs again, until the model emits a stop token or hits a length limit. Training uses the same boxes with one addition, a loss, at the end (section 6).

Chapter 1

Text to ids: the coat check

Try the coat check first: hand over some text and see what comes back.

Everyday picture A coat check. You hand over a coat (a piece of text) and get back a numbered ticket (a token id). The model only ever handles the tickets.

Tiny worked example With this module's toy tokenizer, "Reset your password" becomes 5 tickets: Re set y our password → [346, 377, 309, 328, 291]. The common word " password" got a single ticket; the rarer pieces got several.

Figure 2 · Diagram

Reading it: one box, text in and integers out. Everything to the right of this box works only with integers and vectors.

The code tok.encode(prompt); how the kit is learned is the whole of primer.ml.tokenization.

In code: build_pipeline trains the toy tokenizer (primer.ml.tokenization.ByteBPE) and builds an untrained primer.ml.transformer.TinyGPT sized to its vocabulary.

Why it matters Prompt length, price and context limits are all counted in these tickets.

Chapter 2

Ids to vectors: a lookup, not a computation

Everyday picture A dictionary where ticket number 291 opens to page 291, and each page holds a list of numbers describing that token. A second dictionary, indexed by seat number, describes where the token sits.

Tiny worked example The table has 400 rows (one per token) of 32 numbers. Id 291 fetches row 291. The 5 ids fetch 5 rows: a 5 × 32 grid. Row i of the position table (i = 0 to 4) is added to row i of that grid.

Figure 3 · Diagram

Reading it: two lookups, one add. Nothing is multiplied here, rows are simply fetched, which is why this step is nearly free. The tables themselves are learned during training, so similar tokens end up with similar rows (see primer.ml.embeddings).

The math and the code

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
a token's position in the prompt, from 0 4 (the last token)
the token id at position = 291
the token embedding table 400 × 32
row of that table row 291: 32 numbers
row of the position table row 4: 32 numbers
what the first block receives for position 32 numbers

In words: "each token's input vector is its token's row plus its position's row."

With the numbers: , one 32-number list plus another. The code is model.wte[ids] + model.wpe[:len(ids)]. In miniature, with 3-number rows instead of 32: = (0.2, −0.1, 0.5) and = (0.1, 0.3, −0.2) give = (0.3, 0.2, 0.3).

Level 3: in Python
# row 291 of the token table (3 numbers, not 32)
E_291 = [0.2, -0.1, 0.5]
# row 4 of the position table
P_4 = [0.1, 0.3, -0.2]
# E_(t_i) + P_i, number by number
x_4 = [round(e + p, 2) for e, p in zip(E_291, P_4)]
x_4  # → [0.3, 0.2, 0.3]

In code: trace does both lookups and the add, and keeps every stage's result so you can inspect the grid before and after positions are mixed in.

Why it matters This is the only place a token's identity enters the model; every later step works on these vectors.

Now every ticket of the prompt at once: fetch what it is, fetch where it sits, add.

Chapter 3

The transformer blocks: rounds of meeting and desk work

Everyday picture The team from primer.ml.transformer: each round is a meeting where every token listens to the others (attention), then desk work where each token thinks alone (feed-forward). This toy runs 2 rounds; large models run dozens.

Tiny worked example The 5 × 32 grid goes into block 1 and comes out 5 × 32; the same through block 2. By the end, the vector at the last position ("password") has absorbed information from "Re", "set", "y" and "our".

Figure 4 · Diagram

Reading it: the shape never changes, only the contents. Each block edits every token's vector by adding what it learned from the others.

The math and the code

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the 5 × 32 input grid from section 2
the -th transformer block = 2
"and so on, for every block in between"
the final layer norm
the context-aware vectors 5 × 32

In words: "run the input through every block in turn, then normalize."

With the numbers: here N = 2, so h = LN(Block₂(Block₁(x))). The code is the for block in model.blocks loop in trace. In miniature, with one 3-number vector and two stand-in blocks that each add an edit: x = (1, 2, 3), Block₁ adds (0, 1, 0) and Block₂ adds (0, 0, 2), giving (1, 3, 5). The final norm subtracts the mean (3) and divides by the spread (1.63), so h = (−1.22, 0, 1.22).

Level 3: in Python
import statistics
x = [1.0, 2.0, 3.0]
# a stand-in edit
def block_1(v): return [v_i + d for v_i, d in zip(v, [0, 1, 0])]
def block_2(v): return [v_i + d for v_i, d in zip(v, [0, 0, 2])]
def LN(v):
    mu, sigma = statistics.fmean(v), statistics.pstdev(v)
    # centre, then rescale
    return [round((v_i - mu) / sigma, 2) for v_i in v]
h = x
# Block_1 first, then Block_2 ... up to Block_N
for block in [block_1, block_2]:
    h = block(h)
h  # → [1.0, 3.0, 5.0]
LN(h)  # → [-1.22, 0.0, 1.22]

Why it matters This is where nearly all the compute and all the "understanding" happen.

Chapter 4

Scores, softmax and temperature

Turn the temperature yourself before reading about it.

Everyday picture A scoreboard with one line for every token in the vocabulary. Softmax turns the scores into shares of a pie. Temperature sets how adventurous the pick is: low temperature almost always takes the biggest slice; high temperature gives the small slices a real chance.

Tiny worked example Three candidate tokens scored 2.0, 1.0 and 0.5.

temperature p(A) p(B) p(C)
0 (greedy) 1.000 0.000 0.000
0.5 0.844 0.114 0.042
1 0.629 0.231 0.140
2 0.481 0.292 0.227

Figure 5 · Diagram

Reading it: only the last position's vector is used to choose the next token. It is scored against every token's embedding, the scores are divided by the temperature, softmax turns them into probabilities that sum to 1, and one token is drawn at random in proportion to them.

The math and the code Softmax raises e (≈ 2.718) to the power of each score, so bigger scores get disproportionately bigger shares, then divides by the total so the shares add up to 1:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the score of candidate token = 2.0
the temperature 0.5
e ≈ 2.718 raised to that power = 54.6
number of candidates (the vocabulary size) 3 here, 400 in the model
add up over every candidate
probability that token comes next 0.844

In words: "the chance of token i is e to the power of its score over the temperature, divided by the same quantity summed over all tokens."

With the numbers: T = 0.5 doubles every score to (4, 2, 1): e⁴ = 54.6, e² = 7.39, e¹ = 2.72, total 64.7, so p(A) = 54.6 / 64.7 = 0.844 (next_token_probs). Temperature 0 is the limit case: all probability on the top score.

Level 3: in Python
import math
z = [2.0, 1.0, 0.5]
T = 0.5
# e^(z_i / T) for each candidate
exps = [math.exp(z_i / T) for z_i in z]
[round(e, 2) for e in exps]  # → [54.6, 7.39, 2.72]
# Σ_j e^(z_j / T)
total = sum(exps)
round(total, 1)  # → 64.7
# p_i: each share of the total
[round(e / total, 3) for e in exps]  # → [0.844, 0.114, 0.042]

Figure 6 · Chart

token A (score 2.0) token B (1.0) token C (0.5) 0.0 0.2 0.4 0.6 0.8 1.0 probability Temperature: low sharpens, high flattens T = 0.25 T = 0.5 T = 1.0 T = 2.0

At T = 0.25 token A takes 98% of the probability; at T = 2 the three tokens share it 48%, 29% and 23%

Reading it: three groups of bars, one per candidate token; within each group the bars run from low temperature (left) to high (right). For token A the bars fall as temperature rises, and for tokens B and C they rise. At T = 0.25, A takes almost everything; at T = 2 the three are much closer.

Figure 7 · Chart

'b' 'q' '3' ' s' ' i' '�' 'set' 'N' ' b' 0.0000 0.0005 0.0010 0.0015 0.0020 0.0025 0.0030 0.0035 probability of being next Untrained model: 'The cat sat on the' → ? uniform guess 1/400

The untrained model's 12 favourites are random byte fragments, each only 1.25 to 1.4 times the uniform 1-in-400 share

Reading it: the bars are the 12 most probable next tokens after "The cat sat on the", according to our untrained model; the dashed line is a uniform guess, 1 in 400 (0.25%). Even these favourites clear the line only modestly: the top one gets about 0.35%, 1.4 times the uniform share, and the twelfth about 1.25 times. Across all 400 tokens every probability stays between 0.74 and 1.39 times uniform, and the "favourites" are random byte fragments. That is exactly what random weights should give: a nearly flat guess with small random bumps. The machinery works, but nothing has been learned yet.

Why it matters Use low temperature for extraction and tool calls, where you want the most likely answer, and higher temperature for creative writing. Temperature 0 reduces randomness but doesn't guarantee identical outputs on real serving hardware.

Chapter 5

The loop: append and repeat

Everyday picture Writing a sentence one word at a time, and re-reading everything you've written before choosing each new word.

Tiny worked example A 2-token prompt, 4 new tokens. Step 1 reads 2 tokens, step 2 reads 3, step 3 reads 4, step 4 reads 5: 14 token positions to produce 4 tokens. With a 20-token prompt and 200 new tokens, the naive loop reads 23,900 positions; a KV cache reads 219.

Figure 8 · Diagram

Reading it: the loop box is the whole of generation. The expensive arrow is "full forward pass over ALL ids": every step re-reads text that hasn't changed. Because of the causal mask, earlier tokens' internal vectors can't change when new tokens arrive, so that repeated work is pure waste. The KV cache (primer.ml.inference) stores it once.

The math Total positions processed without a cache:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
prompt length in tokens 2
tokens to generate 4
the step counter, from 0 to 0, 1, 2, 3
tokens re-read at step 2, 3, 4, 5
total positions run through the model 14

In words: "each step re-reads the prompt plus everything generated so far; add that up over all steps."

With the numbers: 4 × 2 + (4 × 3) / 2 = 8 + 6 = 14, the count generate reports as positions_processed.

Level 3: in Python
p, n = 2, 4
# Σ over t = 0 .. n-1 of (p + t): 2 + 3 + 4 + 5
sum(p + t for t in range(n))  # → 14
# the closed form gives the same count
n * p + n * (n - 1) // 2  # → 14

Figure 9 · Chart

0 25 50 75 100 125 150 175 200 tokens generated (after a 20-token prompt) 0 5000 10000 15000 20000 25000 token positions run through the model Why the naive loop is too slow no cache: re-read everything KV cache: one new token per step

Without a cache work grows with the square of output length, 23,900 positions at 200 tokens against 219 with a KV cache

Reading it: the x-axis is how many tokens have been generated after a 20-token prompt; the y-axis is total work in token positions. The red curve bends upward: it grows with the square of the output length. The blue line (with a KV cache) grows by exactly one per token. At 200 tokens the gap is more than a hundredfold.

In code: Generation holds the generated ids, the decoded text and the positions count; naive_vs_cached_work counts the positions processed with and without a KV cache for the figure.

Why it matters Output length drives latency and cost; this is why every serving system caches keys and values.

Chapter 6

Training: the same forward pass, plus a loss

Everyday picture A guessing game with instant feedback. Cover the next word, guess it, uncover it, and note how surprised you were. Training nudges every weight to make the surprise smaller next time.

Tiny worked example If the model gives the right next token probability 0.5, the loss is −ln 0.5 = 0.693. Our untrained model averages 5.97 on a real sentence, close to ln 400 = 5.99, the score of a blind uniform guess.

Figure 10 · Diagram

Reading it: the first two boxes are the forward pass you just traced. Training adds the loss (how surprised the model was by the real next tokens), then backpropagation (primer.ml.neural_net) works out how each weight contributed, and the optimizer (primer.ml.optimizers) adjusts them. One pass over n tokens gives n − 1 guesses at once, in parallel, which is a big reason transformers train fast.

The math and the code A logarithm answers "to what power must I raise e to get this number?" For probabilities between 0 and 1 it is negative, so we flip the sign; ln 1 = 0 (no surprise) and ln of a tiny number is very negative (huge surprise).

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
tokens in the training text 12 ("The model reads tokens, not words.")
the token at position
probability the model gave the actual next token, having seen everything before it; the bar reads "given" ≈ 1/400 when untrained
natural logarithm ln(1/400) = −5.99
the average over all guesses
the loss training pushes down 5.97

In words: "for every position, take the log of the probability the model gave the true next token, average them, and flip the sign."

With the numbers: a uniform guess over 400 tokens gives every true token p = 1/400, so each term is −ln(1/400) = ln 400 = 5.99. Our untrained model scores 5.97, so it is still essentially guessing. Perplexity, e raised to the loss, is 393: "as unsure as choosing among 393 equally likely tokens" (next_token_loss, cross_entropy).

Level 3: in Python
import math
# one guess that gave the true token p = 0.5
round(-math.log(0.5), 3)  # → 0.693
n = 12
# a uniform guess gives every true token 1/400
p = [1 / 400] * (n - 1)
# -(1/(n-1)) Σ ln p
L = -sum(math.log(p_i) for p_i in p) / (n - 1)
round(L, 2)  # → 5.99

Why it matters Pretraining is exactly this, over trillions of tokens: predicting the next token well forces the model to absorb grammar, facts and reasoning patterns. Random weights produce gibberish, and training is what turns the same machinery into a useful model.

The loss one guess at a time: how surprised the model was, and the average that training pushes down.

Test yourself

6 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1What happens, step by step, when you send a prompt?Think it through, then reveal

Tokenizer turns text into ids; each id looks up an embedding; position information is added; the vectors pass through N transformer blocks; the last position's vector is scored against the whole vocabulary; softmax and sampling pick a token; it's appended and the loop repeats until a stop token.

Question 2Why does only the last position matter when generating?Think it through, then reveal

Its vector has attended to every earlier token and is the one trained to predict what comes next; earlier rows predict tokens we already have.

Question 3What does temperature do to the scores (2, 1, 0.5) at T = 0.5?Think it through, then reveal

Doubles them to (4, 2, 1) before softmax, sharpening the distribution: the top token rises from 0.63 to 0.84.

Question 4How many token positions does the naive loop process for a 2-token prompt and 4 new tokens?Think it through, then reveal

2 + 3 + 4 + 5 = 14. A KV cache avoids re-reading the unchanged prefix.

Question 5What loss does an untrained model get over a 400-token vocabulary, and why?Think it through, then reveal

About ln 400 ≈ 5.99 (perplexity about 400), because with random, tiny weights its predictions are close to uniform.

Question 6How is training different from inference?Think it through, then reveal

Same forward pass, plus a loss comparing predictions to the real next tokens, then backpropagation and a weight update. Inference only runs the forward pass.

Primary sources

The papers behind this lesson

Vaswani et al. (2017), Attention Is All You Need.

The architecture inside the "transformer blocks" box.

Read on rumblr →The paper ↗
Radford et al. (2019), Language Models are Unsupervised Multitask Learners (GPT-2).

The decoder-only, next-token-prediction recipe this pipeline follows, including learned positions, byte-level BPE and tied embeddings.

The paper ↗
Brown et al. (2020), Language Models are Few-Shot Learners (GPT-3).

Showed that scaling this same loop up produces models that follow instructions from examples in the prompt.

Read the annotated companion →The paper ↗
Holtzman et al. (2019), The Curious Case of Neural Text Degeneration.

Introduced nucleus (top-p) sampling and explained why pure greedy decoding produces repetitive text.

The paper ↗

Researcher's shelf

Further reading

  • Andrej Karpathy, Let's build GPT (video): https://www.youtube.com/watch?v=kCc8FmEb1nY
  • Karpathy's nanoGPT: https://github.com/karpathy/nanoGPT
  • Jay Alammar, The Illustrated GPT-2: https://jalammar.github.io/illustrated-gpt2/
  • 3Blue1Brown, But what is a GPT? (video): https://www.youtube.com/watch?v=wjZofJX0v4M
  • Hugging Face, How to generate text (decoding strategies): https://huggingface.co/blog/how-to-generate

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.