At a glance
Key takeaways
- Features are directions in activation space, not individual neurons; a dot product with the right direction reads a feature out.
- Probes are small linear classifiers trained on frozen activations. They show information is present, not that the model uses it; score them on unseen data and against a control task.
- The logit lens applies the model's own output layer to intermediate residual states, showing what it would predict after each layer.
- Activation patching copies one activation from a clean run into a corrupted run and measures how much of the right answer returns: the basic causal experiment.
- Superposition: when features are sparse, models store more features than they have neurons, as nearly perpendicular directions, which makes neurons polysemantic.
- Sparse autoencoders learn an overcomplete dictionary with an L1 penalty so each input uses a few latents, recovering features from superposition, at the cost of some unexplained variance.
Level 2
How it works, from scratch
Everyday picture A car makes a strange noise. You can take it for a test drive and note when the noise happens: that is testing the car's behaviour. Or you can open the hood, put a stethoscope on the engine, and swap parts until the noise stops: that is looking at the mechanism. Test drives tell you what the car does. Only the open hood tells you why.
Every other lesson in this primer treats a trained model as something to
build, train or call. Evaluations (primer.agents.evals) are test drives:
they measure what the model says. Interpretability opens the hood. It
asks what the model's billions of internal numbers represent, and which of
them cause the answer. Three reasons to care:
- Debugging. When a model gets something wrong, you want to know which step failed, the same way you read a stack trace rather than just the error message.
- Trust. A model can be right for the wrong reason. An image classifier that "detects" wolves by looking for snow in the background scores well until it meets a wolf on grass.
- Safety. The output is only part of what a model computes. Some questions (is it relying on a stereotype? does it know more than it says?) can only be answered by reading the computation itself.
A tiny worked example. This lesson asks four questions of one kind of sentence, "the Eiffel Tower is in …" → "Paris", each with its own tool, on toy models small enough to check by hand:
| Question | Tool | What our toy shows |
|---|---|---|
| Is a property stored in this hidden state? | a probe | sentiment, read at 99% on unseen examples |
| What would the model say if it stopped at layer ℓ? | the logit lens | "London", then "Paris", then "Paris" |
| Which activations cause the answer? | activation patching | the subject word early, the last word late |
| What are the model's units of meaning? | superposition and sparse autoencoders | five features packed into two neurons, then recovered |
Figure 1 · Diagram
flowchart LR M["A trained model<br/>(frozen)"] --> A["Hidden activations<br/>at every layer and word"] A --> P["Probe:<br/>is property X in here?"] A --> L["Logit lens:<br/>what would it predict now?"] A --> AP["Activation patching:<br/>does this activation cause the answer?"] A --> S["Sparse autoencoder:<br/>which features make up this state?"] P & L --> R["Reading<br/>(correlation)"] AP --> C["Intervening<br/>(causation)"] S --> U["Units of analysis<br/>(what to read and intervene on)"]
Chapter 1
A feature is a direction
Everyday picture A band with two instruments plays through two speakers. The sound engineer pans the guitar mostly to the right speaker and the piano mostly to the left, but each speaker plays a mix of both. If you unplug one speaker you don't lose "the guitar"; you lose part of everything. The instruments are not the speakers. Each instrument is a setting across all speakers, a direction, and the sound in the room is the sum of every instrument times how loudly it plays.
A model's hidden state works the same way. The speakers are neurons: the individual numbers in a layer's output vector. The instruments are features: the things the model has learned to track, such as "this review is positive" or "this verb is in the past tense". Each feature is stored as a direction in the space of neurons, and the hidden state is the sum of every active feature's direction, scaled by how strongly it is present.
A tiny worked example. Take a hidden layer with two neurons and two features whose directions are "positive" = (0.6, 0.8) and "past tense" = (0.8, −0.6). A sentence that is quite positive (0.9) and a little past-tense (0.4) has the hidden state
0.9 · (0.6, 0.8) + 0.4 · (0.8, −0.6) = (0.54 + 0.32, 0.72 − 0.24) = (0.86, 0.48).
Neuron 1 reads 0.86 and neuron 2 reads 0.48. Neither number is "how
positive" or "how past-tense": each neuron is a blend of both features.
To get the features back, take the dot product (multiply matching
entries and add; see primer.notation) of the hidden state with each
direction:
- positive: 0.6 · 0.86 + 0.8 · 0.48 = 0.516 + 0.384 = 0.9
- past tense: 0.8 · 0.86 − 0.6 · 0.48 = 0.688 − 0.288 = 0.4
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the hidden state: one number per neuron | (0.86, 0.48) | |
| how many features are present | 2 | |
| which feature | 1 = positive, 2 = past tense | |
| how strongly feature is present | , | |
| feature 's direction: a list with one number per neuron, of length 1 | ||
| add up one term per feature | two terms | |
| the amount of feature we read back (the hat means "estimated") | 0.9 | |
| dot product: multiply matching entries, then add |
In words: "the hidden state is the sum of each feature's direction times its strength; to read a feature back, dot the hidden state with that feature's direction."
With the numbers: $h = 0.9 \cdot (0.6, 0.8) + 0.4 \cdot (0.8, -0.6) = (0.86, 0.48)\hat{f}_1 = (0.6, 0.8) \cdot (0.86, 0.48) = 0.9$. The read-back is exact here because the two directions are perpendicular (their dot product is 0.6 · 0.8 + 0.8 · (−0.6) = 0) and each has length 1. Hold on to that condition: the section on superposition is about what happens when it fails.
Level 3: in Python
positive = [0.6, 0.8]
past = [0.8, -0.6]
# h = Σ f_i d_i
h = [0.9 * p + 0.4 * q for p, q in zip(positive, past)]
[round(x, 2) for x in h] # → [0.86, 0.48]
# f̂_i = d_i · h
round(sum(d * x for d, x in zip(positive, h)), 2) # → 0.9
round(sum(d * x for d, x in zip(past, h)), 2) # → 0.4
# the directions are perpendicular: their dot product is 0
round(sum(p * q for p, q in zip(positive, past)), 2) # → 0.0
Figure 2 · Drawn from the lesson's code
Two perpendicular arrows for the positive and past-tense directions, and the hidden state (0.86, 0.48) as their weighted sum, with dotted lines dropping onto each arrow at 0.9 and 0.4
That features are directions, and that most of what a model tracks can be
read with a dot product, is called the linear representation
hypothesis. It is a hypothesis, not a law, but it holds often enough to
power everything below. You have met it before: word-vector arithmetic
such as king − man + woman ≈ queen (primer.ml.embeddings.word2vec) works
because "royalty" and "gender" are directions.
In code: compose_features builds a hidden state from feature amounts and read_features reads them back with dot products.
Chapter 2
Probes: can a straight line read it out?
Everyday picture A doctor can't ask your liver how it is doing, but a blood test can measure a marker that tells them. A probe is a blood test for a hidden state: a tiny classifier, trained by you, that looks at a layer's activations and answers one yes/no question, such as "is this review positive?". The model itself is frozen; only the probe learns.
A tiny worked example. A probe for "positive" on a two-neuron layer has weights w = (1, 0.5) and bias b = 0. On the hidden state h = (2, −1):
- Score it: 1 · 2 + 0.5 · (−1) + 0 = 1.5.
- Squash the score into a probability with the sigmoid σ(z) = 1 / (1 + e^−z), which maps any number into the range 0 to 1: σ(1.5) = 1 / (1 + 0.223) = 0.818.
The probe is 82% sure this hidden state belongs to a positive review. On h = (−1, 1) the score is −1 + 0.5 = −0.5 and σ(−0.5) = 0.378: probably negative.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the frozen model's hidden state for one example | (2, −1) | |
| the probe's weights: one per neuron, learned | (1, 0.5) | |
| the probe's bias: one number, learned | 0 | |
| the probe's raw score (its logit) | 1.5 | |
| the sigmoid: turns any score into a probability between 0 and 1 | ||
| Euler's number, about 2.718 | ||
| the probe's probability that the property is present | 0.818 |
In words: "dot the hidden state with the probe's weights, add the bias, and squash the result into a probability."
With the numbers: $p = \sigma(1 \cdot 2 + 0.5 \cdot (-1) + 0) = \sigma(1.5) = 1 / (1 + 0.223) = 0.818$.
Level 3: in Python
import math
w = [1.0, 0.5]
b = 0.0
h = [2.0, -1.0]
# w · h + b
z = sum(wi * hi for wi, hi in zip(w, h)) + b
z # → 1.5
# σ(z)
round(1 / (1 + math.exp(-z)), 3) # → 0.818
# a second hidden state, (−1, 1), reads as probably negative
round(1 / (1 + math.exp(-(-1 * 1.0 + 1 * 0.5))), 3) # → 0.378
This is exactly logistic regression (primer.ml.neural_net), trained
the usual way: show it examples whose answer you know, measure its
binary cross-entropy (primer.ml.losses), and nudge w and b downhill. The
gradient of that loss with respect to the score is simply (p − y), the gap
between the probe's probability and the true 0/1 label, which makes each
update one line of code.
Figure 3 · Diagram
flowchart LR T["Text with a known label<br/>(positive or negative)"] --> M["Frozen model<br/>(no weights change)"] M --> H["Hidden state h<br/>at the layer under study"] H --> P["Probe: σ(w · h + b)"] P --> Y["Probability the label is 'positive'"] Y --> G["Compare with the true label<br/>(cross-entropy)"] G -. "gradient updates w and b only" .-> P
On our toy model. PlantedModel stores two known features in a
32-neuron hidden layer: sentiment and tense, each as ±1 along its own
random direction, plus random noise of size 0.4 on every neuron. No single
neuron holds either feature. A probe trained on just 100 examples reads
sentiment correctly on 99% of 1,000 examples it never saw.
Why keep probes simple, and check them against a control. A probe can succeed for the wrong reason. Give a probe random labels, a control task with nothing real to find, and it still fits 73% of its 100 training examples, because 33 adjustable numbers can memorize a lot of noise. On unseen examples it scores 50%, a coin flip. Two habits follow: always score a probe on examples it never saw, and compare it with a control task, so that "the probe works" means "the hidden state encodes this" rather than "the probe is clever". This is also why probes are kept linear: a powerful probe (a deep network) could compute the property from raw ingredients by itself, and then its success would tell you about the probe, not the model.
Decodable is not the same as used
Everyday picture A library holds a book nobody ever borrows. Finding it on the shelf proves the library has it, not that it shaped anything any reader did.
A tiny worked example. Our toy model's output reads sentiment and, by construction, gives tense a weight of exactly zero. A probe still reads tense at 98% on unseen examples. Now intervene: take 200 examples, flip only one feature's label, keep everything else fixed, and measure how much the model's output moves.
| Flip | Probe accuracy for it | Output moves by |
|---|---|---|
| sentiment | 99% | 2.00 |
| tense | 98% | 0.00 |
Figure 4 · Drawn from the lesson's code
Left, probe accuracy: sentiment 0.99 and tense 0.98 on unseen examples, random labels 0.73 on training examples but about 0.5 on unseen ones. Right, flipping sentiment moves the output by 2 and flipping tense moves it by 0
A probe finding information does not mean the model uses it. To find out what the model uses you have to change something and watch the output, which is what the rest of this lesson does.
In code: train_probe fits a Probe by gradient descent on frozen hidden states from PlantedModel; probe_report trains the three probes, and flip_effect performs the intervention.
Chapter 3
The logit lens: reading the model's mind mid-thought
Everyday picture A writer keeps every draft of an article. Draft 1 says the landmark is "in a European capital", draft 2 says "probably Paris", the final says "Paris". Reading the drafts shows when the writer made up their mind. The logit lens reads a language model's drafts.
It works because of how a transformer is built (primer.ml.transformer).
Each word has a running vector, the residual stream. Every layer reads
it and adds a correction to it, rather than replacing it. At the very
end, the output layer (the unembedding) turns the final vector into
one score per word in the vocabulary. Because every layer writes into the
same stream, you can apply that final step early, after any layer, and ask
"what would the model predict if it stopped here?"
A tiny worked example. A vocabulary of three words, Paris, London and Rome, a residual stream of two numbers, and an unembedding that scores Paris = first number, London = second number, Rome = minus the first:
| After | Residual | Scores (Paris, London, Rome) | Probabilities | Top guess |
|---|---|---|---|---|
| the embedding | (0.2, 0.3) | (0.2, 0.3, −0.2) | (0.36, 0.40, 0.24) | London |
| layer 1, which adds (0.6, 0) | (0.8, 0.3) | (0.8, 0.3, −0.8) | (0.55, 0.34, 0.11) | Paris |
| layer 2, which adds (1.2, −0.3) | (2.0, 0.0) | (2.0, 0.0, −2.0) | (0.87, 0.12, 0.02) | Paris |
The first draft is a vague "some capital" that happens to lean London; layer 1 tips it to Paris; layer 2 commits.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| which layer we stop after (0 = straight after the embedding) | 2 | |
| the residual stream after layer , for one word | (2.0, 0.0) | |
the model's final normalization (primer.ml.transformer); our 2-number toy has none, so here LN(h) = h |
(2.0, 0.0) | |
| the unembedding: one row per vocabulary word, dotted with the state to give that word's score | rows (1, 0), (0, 1), (−1, 0) | |
| the scores (logits), one per word | (2, 0, −2) | |
| softmax | turns scores into probabilities that add up to 1 (primer.ml.attention) |
(0.87, 0.12, 0.02) |
| what the model would predict if it stopped after layer | Paris at 0.87 |
In words: "take the residual stream partway up, push it through the model's own final normalization and output layer, and read off the probabilities."
With the numbers: after layer 2, $W_U h_2 = (1 \cdot 2 + 0 \cdot 0,; 0 \cdot 2 + 1 \cdot 0,; -1 \cdot 2 + 0 \cdot 0) = (2, 0, -2)$; then , , , total 8.52, so P(Paris) = 7.39 / 8.52 = 0.87.
Level 3: in Python
import math
# rows: Paris, London, Rome
W_U = [[1, 0], [0, 1], [-1, 0]]
h = [0.2, 0.3]
writes = [[0.6, 0.0], [1.2, -0.3]]
# a residual layer adds its write to the stream
for w in writes:
h = [a + b for a, b in zip(h, w)]
[round(x, 2) for x in h] # → [2.0, 0.0]
# W_U h: one score per word
scores = [sum(u * x for u, x in zip(row, h)) for row in W_U]
[round(s, 2) for s in scores] # → [2.0, 0.0, -2.0]
# softmax
exps = [math.exp(s) for s in scores]
[round(e / sum(exps), 2) for e in exps] # → [0.87, 0.12, 0.02]
Figure 5 · Diagram
flowchart LR
E["Embedding"] --> H0(("h₀")) --> L1["Layer 1<br/>adds its write"] --> H1(("h₁")) --> L2["Layer 2<br/>adds its write"] --> H2(("h₂")) --> OUT["Final norm + W_U<br/>(the real output)"]
H0 -.-> LENS0["lens: norm + W_U<br/>London 0.40"]
H1 -.-> LENS1["lens: norm + W_U<br/>Paris 0.55"]
H2 -.-> LENS2["lens: norm + W_U<br/>Paris 0.87"]
primer.ml.transformer.TinyGPT, they do).On a model where we know the answer. LandmarkModel is a two-layer,
one-head transformer built by hand to complete "Eiffel is in" with
"Paris". Layer 1 is an MLP that looks up a fact at every word (landmark
in, city out); layer 2 is an attention head that lets the last word, "in",
find the landmark and copy its city.
Figure 6 · Diagram
flowchart LR T["Eiffel · is · in"] --> EMB["Embed each word"] EMB --> MLP["Layer 1: MLP at every word<br/>Eiffel → writes 'Paris'"] MLP --> ATT["Layer 2: attention head<br/>'in' looks for the landmark,<br/>copies its city"] ATT --> U["Output at 'in':<br/>Paris 3, Rome 0, London 0"]
Figure 7 · Drawn from the lesson's code
Left, the worked example's probabilities for Paris, London and Rome after the embedding, layer 1 and layer 2. Right, a 3 by 3 grid of the lens's probability of Paris by layer and word: the Eiffel column turns dark after layer 1, the in column only after layer 2
The logit lens was first described for GPT-2 in 2020, where middle layers already "guess" the next word surprisingly well. In some models the early layers read as nonsense through the final output layer, because they do not yet speak its "language"; the tuned lens fixes this by training a small translator for each layer before applying the output layer.
In code: logit_lens applies an output layer to any residual states, worked_lens runs the table above, residual_stream and tinygpt_logit_lens do the same for primer.ml.transformer.TinyGPT, and lens_map builds the grid for LandmarkModel.
Chapter 4
Activation patching: which activations cause the answer?
Everyday picture Two cars of the same model sit side by side: one starts, one doesn't. You move parts from the good car into the bad one, one part at a time, and try the ignition after each swap. The part whose swap makes the bad car start is the part that mattered. Activation patching (also called causal tracing) does this with a model's internal activations.
A tiny worked example. Run the landmark model twice:
- the clean prompt "Eiffel is in": scores Paris 3, Rome 0, so the logit difference Paris − Rome is +3;
- the corrupted prompt "Colosseum is in": Paris 0, Rome 3, so the difference is −3.
Now run the corrupted prompt again, but at one chosen place overwrite the activation with the one from the clean run, and see how much of the gap between −3 and +3 comes back:
- patch the MLP output at the landmark's position: the difference jumps back to +3, so 100% of the answer is restored;
- patch the state at "is": it stays at −3, 0% restored ("is" is the same word in both prompts, so its state carries nothing about the landmark).
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| logit difference: the correct answer's score minus the wrong answer's score (Paris − Rome) | ||
| on the clean prompt, nothing patched | +3 | |
| on the corrupted prompt, nothing patched | −3 | |
| on the corrupted prompt, with one clean activation copied in | +3, −3, or anything between | |
| restored | the fraction of the gap that one patch closes: 0 = nothing, 1 = everything | 1, 0 |
In words: "how far the patch moved the answer, as a share of the full distance from the corrupted answer to the clean one."
With the numbers: patching the landmark's MLP output gives (3 − (−3)) / (3 − (−3)) = 6 / 6 = 1. Patching "is" gives (−3 − (−3)) / 6 = 0. A patch that left the difference at 0 would give (0 − (−3)) / 6 = 0.5: half the answer restored.
Level 3: in Python
clean, corrupt = 3.0, -3.0
# restored = (LD_patched − LD_corrupt) / (LD_clean − LD_corrupt)
def restored(patched):
return (patched - corrupt) / (clean - corrupt)
# the landmark's MLP output, patched in
restored(3.0) # → 1.0
# the state at "is", patched in
restored(-3.0) # → 0.0
# a patch that only reaches a tie
restored(0.0) # → 0.5
Why a difference of two scores, rather than the probability of Paris? Because the corrupted prompt differs from the clean one in exactly one fact, the difference measures exactly that fact, and it moves smoothly, where a probability can saturate near 0 or 1 and hide a change.
Figure 8 · Diagram
flowchart TB
subgraph C["1. Clean run: 'Eiffel is in'"]
c1["save every activation"] --> c2["Paris − Rome = +3"]
end
subgraph K["2. Corrupted run: 'Colosseum is in'"]
k1["Paris − Rome = −3"]
end
subgraph P["3. Patched run: corrupted prompt,<br/>ONE activation taken from the clean run"]
p1["Paris − Rome = ?"]
end
c1 -- "copy one activation" --> P
P --> R["fraction restored =<br/>(? − (−3)) / (3 − (−3))"]
K --> R
C --> R
Figure 9 · Drawn from the lesson's code
A 3 by 3 grid, layers by positions, of the fraction of the answer restored by one patch: 1 at the Eiffel column for the embedding and layer 1, 1 at the in column after layer 2, 0 everywhere else
Compare with the lens grid above. After layer 2, the lens reads "Paris" at the Eiffel position with probability 0.995, yet patching that state restores 0.00: after the last layer nothing reads that position any more. The information is there, and it no longer matters. Readable is not the same as used, again.
Why it matters in practice. This is how researchers located where GPT-style models store facts: causal tracing on prompts like ours found factual recall concentrated in middle-layer MLPs at the subject's last token, and then used that to edit single facts. The same method, one site at a time, has traced whole circuits, such as the one GPT-2 small uses to fill in "When Mary and John went to the store, John gave a drink to" → "Mary".
In code: LandmarkModel is the hand-built model and LandmarkModel.run accepts patches; fraction_restored is the formula and patching_map patches every layer and position in turn.
Chapter 5
Superposition: more features than neurons
Everyday picture Back to the band, but now five instruments play through two speakers. The engineer pans each instrument to its own position around the room: hard left, front right, back left, and so on. When one instrument plays alone, you can tell which one by where the sound comes from. When two play at once, the positions blur together and you might mistake the pair for a third instrument. The trick works because in this band, most of the time, only one instrument is playing.
Models face the same squeeze. There are far more concepts in the world than neurons in a layer, but in any one sentence almost all of them are absent: features are sparse. So models store more features than they have dimensions, as directions that are nearly perpendicular rather than exactly. That is superposition.
A tiny worked example. Put five features in two neurons, spread 72° apart like the points of a pentagon. Feature 's direction is (cos 72k°, sin 72k°): feature 0 is (1, 0), feature 1 is (0.309, 0.951), and so on. Switch on feature 0 alone, at strength 1. The hidden state is (1, 0). Reading every feature back with a dot product gives
| Feature | Angle from feature 0 | Read-back (cos of the angle) | After adding −0.31 and ReLU |
|---|---|---|---|
| 0 | 0° | 1.000 | 0.69 |
| 1 | 72° | 0.309 | 0 |
| 2 | 144° | −0.809 | 0 |
| 3 | 216° | −0.809 | 0 |
| 4 | 288° | 0.309 | 0 |
Five directions can't all be perpendicular in two dimensions, so reading feature 0 leaks 0.309 into each neighbour: interference. The fix is a small negative bias and a ReLU (which turns negatives into zero): the leak of 0.309 falls below the bias of 0.31 and is filtered to 0, and the real feature survives at 0.69. The cost comes when two neighbours are on at once: each then reads 1 + 0.309 − 0.31 = 0.999 instead of 0.69, too high. Sparsity is a bet that such collisions are rare.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the true features: one strength per feature | (1, 0, 0, 0, 0) | |
| the squeeze: 2 rows (neurons) by 5 columns (features); column is feature 's direction | the pentagon | |
| the hidden state: 2 numbers holding 5 features | (1, 0) | |
| transposed (rows and columns swapped): reads each feature back with a dot product | ||
| one bias per feature, learned; negative values filter small leaks | −0.31 each | |
| ReLU | keep positives, turn negatives into 0 | |
| the features read back out | (0.69, 0, 0, 0, 0) | |
| feature 's direction dotted with itself (its squared length) | 1 | |
| how much feature leaks into feature : 0 if perpendicular | 0.309 for neighbours | |
| add up over every other feature |
In words: "squeeze the features into a few neurons, read each one back by dotting with its direction, subtract a threshold and drop anything negative. What you read for feature is its own strength, plus a leak from every other active feature, minus the threshold."
With the numbers: with feature 0 alone on, $\hat{x}_1 = \text{ReLU}(1 \cdot 0 + 0.309 \cdot 1 - 0.31) = \text{ReLU}(-0.001) = 0$ and . With features 0 and 1 both on, .
Level 3: in Python
import math
# feature k's direction: (cos 72k°, sin 72k°)
W = [[math.cos(math.radians(72 * k)), math.sin(math.radians(72 * k))] for k in range(5)]
def read_back(x, bias, relu=True):
# h = W x: two numbers
h = [sum(W[k][n] * x[k] for k in range(5)) for n in range(2)]
# Wᵀ h + b: one number per feature
z = [W[k][0] * h[0] + W[k][1] * h[1] + bias for k in range(5)]
# ReLU keeps positives and turns negatives into 0
return [round(max(0.0, v) if relu else v, 3) for v in z]
# the leak: feature 0 alone, no bias, no ReLU
read_back([1, 0, 0, 0, 0], bias=0, relu=False) # → [1.0, 0.309, -0.809, -0.809, 0.309]
# the bias and ReLU filter it
read_back([1, 0, 0, 0, 0], bias=-0.31) # → [0.69, 0.0, 0.0, 0.0, 0.0]
# two neighbours on: each reads too high
read_back([1, 1, 0, 0, 0], bias=-0.31) # → [0.999, 0.999, 0.0, 0.0, 0.0]
Figure 10 · Diagram
flowchart LR X["5 features<br/>x₀ … x₄<br/>(mostly zero)"] --> W["squeeze: W<br/>(2 × 5)"] W --> H["hidden state<br/>2 neurons"] H --> WT["read back: Wᵀ<br/>(5 × 2)"] WT --> B["+ bias b<br/>(negative)"] B --> R["ReLU"] R --> XH["5 features<br/>read back"]
Training it. Let the model learn and itself, by gradient descent on the reconstruction error, with features that matter less and less (feature is weighted ). Compare two worlds:
- Dense: every feature is on in every example. The model keeps the two most important features, perpendicular, at length 1.00, and gives the other three length 0: it simply drops them.
- Sparse: each feature is on only 5% of the time. The model keeps all five, 72° apart, at length about 1.1, and learns a negative bias of about −0.23 to filter the leaks.
Figure 11 · Drawn from the lesson's code
Three panels. Left, dense features: two perpendicular arrows and three near-zero ones. Middle, sparse features: five arrows spread evenly around a pentagon. Right, bars of how much each of five features moves neuron 1 and neuron 2
That last point is why looking at single neurons so often fails. A neuron that responds to several unrelated features is polysemantic. Vision researchers found neurons that respond to cat faces and to the fronts of cars; language models are full of neurons like that. Superposition explains why: when features outnumber neurons, the features cannot line up with the neurons, so every neuron is a mixture.
In code: pentagon builds the five directions, superposition_readout is the formula, train_superposition learns W and b with gradients from superposition_loss_and_grads (checked against finite_difference_gradient), and features_a_neuron_responds_to lists a neuron's features.
Chapter 6
Sparse autoencoders: getting the features back
Everyday picture A sound engineer receives the two-speaker recording of the five-instrument band, with no notes on who played when. She knows one thing about this band: usually only one instrument plays at a time. Many different scores could produce the same sound, but she writes down the one that uses the fewest instruments. That preference is what lets her recover the real parts rather than some arbitrary mixture.
A sparse autoencoder (SAE) does this for a model's hidden states. It learns a dictionary: many more directions than the layer has neurons (it is overcomplete), trained so that every hidden state can be rebuilt from just a few of them. Each learned direction is a candidate feature, and each one's strength on an input is its latent activation.
A tiny worked example. The hidden state is h = (0.8, 0.6) and the dictionary has three directions, (1, 0), (0, 1) and (0.8, 0.6). Two codes rebuild h perfectly:
| Code (strength of each direction) | Rebuilt | Error | Penalty with λ = 0.1 | Loss |
|---|---|---|---|---|
| A: (0.8, 0.6, 0), two directions | (0.8, 0.6) | 0 | 0.1 × (0.8 + 0.6) = 0.14 | 0.14 |
| B: (0, 0, 1), one direction | (0.8, 0.6) | 0 | 0.1 × 1.0 = 0.10 | 0.10 |
Rebuilding alone can't choose between them. The penalty on the total
strength, the L1 penalty (primer.ml.regularization), prefers the code
that uses one direction. One side effect: the best strength for direction
3 is not quite 1. With strength , the loss is , which
is lowest at (loss 0.0975): the penalty always pulls strengths
a little toward zero, known as shrinkage.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| a hidden state from the model under study | (0.8, 0.6) | |
| how many latents (dictionary directions) the SAE has, usually far more than neurons | 3 | |
| , | the encoder's weights and biases: turn a hidden state into latent strengths | |
| the code: one strength per latent, ReLU keeps them ≥ 0 and mostly exactly 0 | (0, 0, 1) | |
| the decoder: column is latent 's direction, kept at length 1 | (1, 0), (0, 1), (0.8, 0.6) | |
| the decoder bias: the typical hidden state, subtracted before encoding and added back after | (0, 0) | |
| the rebuilt hidden state | (0.8, 0.6) | |
| squared length of the rebuild error: add up the squares of its entries | 0 | |
| the sparsity penalty's strength, chosen by you | 0.1 | |
| the size of latent 's strength | ||
| the loss to minimize: rebuild error plus penalty | 0.10 |
In words: "encode the hidden state into many non-negative strengths, decode it back as a weighted sum of dictionary directions, and pay for both the rebuild error and the total strength used."
With the numbers: for code B, $\hat{h} = 1 \cdot (0.8, 0.6) = (0.8, 0.6)0.1 \times 1 = 0.10L = 0.10$. For code A the penalty is .
Level 3: in Python
dictionary = [[1, 0], [0, 1], [0.8, 0.6]]
h = [0.8, 0.6]
lam = 0.1
def loss(code):
# ĥ = Σ f_i · direction_i
rebuilt = [sum(f * d[n] for f, d in zip(code, dictionary)) for n in range(2)]
# ‖h − ĥ‖² + λ Σ |f_i|
return round(sum((a - b) ** 2 for a, b in zip(h, rebuilt)) + lam * sum(abs(f) for f in code), 4)
loss([0.8, 0.6, 0]) # → 0.14
loss([0, 0, 1.0]) # → 0.1
# shrinkage: a slightly weaker code is cheaper still
loss([0, 0, 0.95]) # → 0.0975
Figure 12 · Diagram
flowchart LR H["hidden state h<br/>(d numbers)"] --> ENC["encoder<br/>ReLU(W_e(h − b_d) + b_e)"] ENC --> F["code f<br/>(m ≫ d numbers,<br/>almost all exactly 0)"] F --> DEC["decoder<br/>W_d f + b_d"] DEC --> HH["rebuilt ĥ"] HH --> LOSS["loss = ‖h − ĥ‖² + λ Σ|f|"] H --> LOSS F --> LOSS
On the superposition toy. Take 20,000 hidden states from the pentagon model (each feature on 5% of the time), train an SAE with 5 latents and λ = 0.3, and compare the learned decoder columns with the five planted directions. Every planted feature is matched by a latent with cosine similarity above 0.99 (1 would be the same direction), and an example with a feature on lights up about 1.07 latents on average. The price is shrinkage: 11% of the variance goes unexplained. With a tiny penalty, λ = 0.01, the SAE rebuilds the data perfectly (0.0% unexplained), yet its worst match to a real feature is only 0.80: any two directions can rebuild a two-dimensional space, so without sparsity pressure nothing forces the dictionary to find the model's actual features.
Figure 13 · Drawn from the lesson's code
Left, grey dots of hidden states lying mostly along five rays, with thick grey planted directions and red learned latent arrows lying on top of them. Right, as the penalty grows from 0.003 to 1, the worst feature match peaks near 0.99 at a penalty of 0.3, while the unexplained variance climbs from 0 to above 0.5
Why it matters in practice. This is the tool that turned interpretability from "a few hand-picked neurons" into something that scales. Anthropic's Towards Monosemanticity (2023) trained SAEs on a small transformer and found thousands of features that each fire for one recognizable thing. Scaling Monosemanticity (2024) did the same inside a production model, Claude 3 Sonnet, and found features for concepts such as the Golden Gate Bridge. Turning that one feature up made the model bring up the bridge in almost every answer, which is the causal check that the feature means what it seems to mean.
In code: sae_objective computes the loss for one proposed code, SparseAutoencoder holds the encoder and decoder with its hand-derived SparseAutoencoder.loss_and_grads, train_sae fits it, and match_features compares learned directions with planted ones.
Chapter 7
What these tools cannot (yet) show
Everyday picture A brain scan shows which regions light up while a person reads, but not the sentence they are thinking. Every tool in this lesson is like that: a real instrument that measures something true and partial.
A tiny worked example. Each of our own toys already showed a limit:
| What we saw | The limit it shows |
|---|---|
| A probe read tense at 98%, and the output ignored tense completely | Probes find information, not use |
| The lens read "Paris" at a position where patching restored 0% | Reading an activation doesn't mean anything downstream reads it |
| The weak SAE rebuilt everything perfectly and still found the wrong directions | A good rebuild score doesn't mean the features are real |
| The strong SAE left 11% of the variance unexplained | Some of the model's activity is in no feature the dictionary found |
| We knew the landmark model's circuit because we built it | In a real model, nobody has the answer key |
Figure 14 · Diagram
flowchart LR A["Correlation<br/>probes, the logit lens:<br/>'this information is here'"] --> B["Intervention<br/>patching, steering:<br/>'changing this changes the output'"] B --> C["Mechanism<br/>a circuit explained end to end:<br/>'this is how the output is computed'"]
The main open problems, in plain words:
- Choices shape answers. Patching results depend on how the corrupted prompt is built (swap one word? add noise?). SAE features depend on the dictionary's size and λ: a bigger dictionary can split one feature into several finer ones.
- Redundancy hides importance. Patching one site at a time can miss parts that back each other up. Models have been observed to partly repair themselves when one component is knocked out, so "patching this restores 0%" does not always mean "this plays no role".
- Coverage. The unexplained part of an SAE's rebuild is not noise to ignore: it is model behaviour no feature describes yet.
- Scale and labour. A circuit that explains one behaviour of a small model can take researchers weeks. No one has a complete account of how a frontier model produces any long answer.
- Labels are human guesses. Naming a feature "Golden Gate Bridge" summarizes the inputs that make it fire; checking that the name is right needs interventions like the one above.
Why it matters in practice. Treat these tools as evidence, stacked
next to behavioural evaluations (primer.agents.evals), never as a
certificate. A probe or a lens is a cheap first look; a patching or
steering experiment is the minimum for a causal claim; and a clean result
on a toy, like every result in this lesson, is where understanding starts,
not where it ends.
Test yourself
6 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1What does it mean to say a feature is a "direction" rather than a neuron?Think it through, then reveal
The feature's presence is stored as a pattern across many neurons: the hidden state moves along a particular direction in proportion to how strongly the feature is present. You read it with a dot product against that direction. Any single neuron is typically a mix of several features.
Question 2A linear probe reads a property from layer 12 at 95% accuracy. What can you conclude, and what can't you?Think it through, then reveal
You can conclude the property is linearly decodable from layer 12, as long as the 95% was on held-out examples and clearly beats a control task with random labels. You can't conclude the model uses it. For that you need an intervention: remove or change that information and see whether the output changes.
Question 3How does the logit lens work, and why does it agree with the model at the last layer?Think it through, then reveal
It takes the residual stream after some layer and applies the model's own final normalization and unembedding, turning it into next-token probabilities. At the last layer that is exactly the computation the model itself performs, so the two must match. Earlier layers are read "as if the model stopped there".
Question 4Describe an activation patching experiment, and what the "fraction restored" means.Think it through, then reveal
Run a clean prompt and save its activations; run a corrupted prompt that changes one fact; then rerun the corrupted prompt with one activation replaced by its clean value. The fraction restored is how much of the clean-minus-corrupted logit difference the patch brings back: 1 means that activation alone carries the fact, 0 means it carries none of it.
Question 5Why do models use superposition, and why does it make neurons hard to interpret?Think it through, then reveal
There are more useful features than neurons. When features are rarely active at the same time, the model can store them as nearly perpendicular directions and filter the small interference with a bias and ReLU. The directions can't line up with the neurons, so each neuron responds to several unrelated features: it is polysemantic.
Question 6Why does a sparse autoencoder need the L1 penalty? What goes wrong if λ is too small or too large?Think it through, then reveal
Many dictionaries rebuild the data equally well; the penalty picks the one where each input uses few latents, which pushes latents onto the real features. Too small and the SAE rebuilds perfectly with meaningless directions; too large and it shrinks activations and leaves much of the model's activity unexplained.
Primary sources
The papers behind this lesson
Introduced linear probes as a way to measure what each layer of a network makes linearly available.
The paper ↗Showed that probes can succeed by memorizing, and introduced control tasks and selectivity to tell the two apart.
The paper ↗Formalized the logit lens and fixed its failures on early layers by training a small translator for each layer.
The paper ↗Introduced causal tracing, found factual recall in middle-layer MLPs at the subject's last token, and edited single facts there.
Read the annotated companion →The paper ↗Used patching to reverse-engineer a complete circuit for one behaviour of a real language model.
The paper ↗Showed with small ReLU models that sparse features are stored in superposition, and when and how the geometry changes.
Read the annotated companion →The paper ↗Trained sparse autoencoders on a small transformer and found thousands of interpretable features hidden in polysemantic neurons.
Read the annotated companion →The paper ↗Scaled sparse autoencoders to a production model and steered its behaviour through the features they found.
Read the annotated companion →The paper ↗Researcher's shelf
Further reading
- Olah et al., Zoom In: An Introduction to Circuits (Distill, 2020): https://distill.pub/2020/circuits/zoom-in/
- Elhage et al., A Mathematical Framework for Transformer Circuits (2021): https://transformer-circuits.pub/2021/framework/index.html
- Elhage et al., Toy Models of Superposition (2022): https://transformer-circuits.pub/2022/toy_model/index.html
- Bricken et al., Towards Monosemanticity (2023): https://transformer-circuits.pub/2023/monosemantic-features/index.html
- Templeton et al., Scaling Monosemanticity (2024): https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html
- Cunningham et al., Sparse Autoencoders Find Highly Interpretable Features in Language Models (2023): https://arxiv.org/abs/2309.08600
- Belinkov, Probing Classifiers: Promises, Shortcomings, and Advances (2021): https://arxiv.org/abs/2102.12452
- Meng et al., Locating and Editing Factual Associations in GPT (2022): https://arxiv.org/abs/2202.05262
- Belrose et al., Eliciting Latent Predictions from Transformers with the Tuned Lens (2023): https://arxiv.org/abs/2303.08112
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.