rumblr Work in progressWIP

● The AI Primer · Lesson 25 · Part 1: how the model works inside

Looking inside the model

This lesson covers Probes, the logit lens, activation patching, superposition, sparse autoencoders

Members · open during launch 52 min14 figures and diagrams
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Features are directions in activation space, not individual neurons; a dot product with the right direction reads a feature out.
  2. Probes are small linear classifiers trained on frozen activations. They show information is present, not that the model uses it; score them on unseen data and against a control task.
  3. The logit lens applies the model's own output layer to intermediate residual states, showing what it would predict after each layer.
  4. Activation patching copies one activation from a clean run into a corrupted run and measures how much of the right answer returns: the basic causal experiment.
  5. Superposition: when features are sparse, models store more features than they have neurons, as nearly perpendicular directions, which makes neurons polysemantic.
  6. Sparse autoencoders learn an overcomplete dictionary with an L1 penalty so each input uses a few latents, recovering features from superposition, at the cost of some unexplained variance.

Level 2

How it works, from scratch

Everyday picture A car makes a strange noise. You can take it for a test drive and note when the noise happens: that is testing the car's behaviour. Or you can open the hood, put a stethoscope on the engine, and swap parts until the noise stops: that is looking at the mechanism. Test drives tell you what the car does. Only the open hood tells you why.

Every other lesson in this primer treats a trained model as something to build, train or call. Evaluations (primer.agents.evals) are test drives: they measure what the model says. Interpretability opens the hood. It asks what the model's billions of internal numbers represent, and which of them cause the answer. Three reasons to care:

  • Debugging. When a model gets something wrong, you want to know which step failed, the same way you read a stack trace rather than just the error message.
  • Trust. A model can be right for the wrong reason. An image classifier that "detects" wolves by looking for snow in the background scores well until it meets a wolf on grass.
  • Safety. The output is only part of what a model computes. Some questions (is it relying on a stereotype? does it know more than it says?) can only be answered by reading the computation itself.

A tiny worked example. This lesson asks four questions of one kind of sentence, "the Eiffel Tower is in …" → "Paris", each with its own tool, on toy models small enough to check by hand:

Question Tool What our toy shows
Is a property stored in this hidden state? a probe sentiment, read at 99% on unseen examples
What would the model say if it stopped at layer ℓ? the logit lens "London", then "Paris", then "Paris"
Which activations cause the answer? activation patching the subject word early, the last word late
What are the model's units of meaning? superposition and sparse autoencoders five features packed into two neurons, then recovered

Figure 1 · Diagram

Reading it: every tool starts from the same place: the numbers a frozen model computes inside itself while it runs. The top two tools only read those numbers, so what they find is a correlation. Patching changes them and watches the output, so what it finds is a cause. Sparse autoencoders answer an earlier question: before you can read or change "a feature", you need to know what the features are. Keep the reading/intervening split in mind; it is the most important distinction in this lesson.

Chapter 1

A feature is a direction

Everyday picture A band with two instruments plays through two speakers. The sound engineer pans the guitar mostly to the right speaker and the piano mostly to the left, but each speaker plays a mix of both. If you unplug one speaker you don't lose "the guitar"; you lose part of everything. The instruments are not the speakers. Each instrument is a setting across all speakers, a direction, and the sound in the room is the sum of every instrument times how loudly it plays.

A model's hidden state works the same way. The speakers are neurons: the individual numbers in a layer's output vector. The instruments are features: the things the model has learned to track, such as "this review is positive" or "this verb is in the past tense". Each feature is stored as a direction in the space of neurons, and the hidden state is the sum of every active feature's direction, scaled by how strongly it is present.

A tiny worked example. Take a hidden layer with two neurons and two features whose directions are "positive" = (0.6, 0.8) and "past tense" = (0.8, −0.6). A sentence that is quite positive (0.9) and a little past-tense (0.4) has the hidden state

0.9 · (0.6, 0.8) + 0.4 · (0.8, −0.6) = (0.54 + 0.32, 0.72 − 0.24) = (0.86, 0.48).

Neuron 1 reads 0.86 and neuron 2 reads 0.48. Neither number is "how positive" or "how past-tense": each neuron is a blend of both features. To get the features back, take the dot product (multiply matching entries and add; see primer.notation) of the hidden state with each direction:

  • positive: 0.6 · 0.86 + 0.8 · 0.48 = 0.516 + 0.384 = 0.9
  • past tense: 0.8 · 0.86 − 0.6 · 0.48 = 0.688 − 0.288 = 0.4
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the hidden state: one number per neuron (0.86, 0.48)
how many features are present 2
which feature 1 = positive, 2 = past tense
how strongly feature is present ,
feature 's direction: a list with one number per neuron, of length 1
add up one term per feature two terms
the amount of feature we read back (the hat means "estimated") 0.9
dot product: multiply matching entries, then add

In words: "the hidden state is the sum of each feature's direction times its strength; to read a feature back, dot the hidden state with that feature's direction."

With the numbers: $h = 0.9 \cdot (0.6, 0.8) + 0.4 \cdot (0.8, -0.6) = (0.86, 0.48)\hat{f}_1 = (0.6, 0.8) \cdot (0.86, 0.48) = 0.9$. The read-back is exact here because the two directions are perpendicular (their dot product is 0.6 · 0.8 + 0.8 · (−0.6) = 0) and each has length 1. Hold on to that condition: the section on superposition is about what happens when it fails.

Level 3: in Python
positive = [0.6, 0.8]
past = [0.8, -0.6]
# h = Σ f_i d_i
h = [0.9 * p + 0.4 * q for p, q in zip(positive, past)]
[round(x, 2) for x in h]  # → [0.86, 0.48]
# f̂_i = d_i · h
round(sum(d * x for d, x in zip(positive, h)), 2)  # → 0.9
round(sum(d * x for d, x in zip(past, h)), 2)  # → 0.4
# the directions are perpendicular: their dot product is 0
round(sum(p * q for p, q in zip(positive, past)), 2)  # → 0.0

Figure 2 · Drawn from the lesson's code

0.0 0.5 1.0 neuron 1 −0.75 −0.50 −0.25 0.00 0.25 0.50 0.75 1.00 neuron 2 Two features, two directions, one hidden state "positive" direction "past tense" direction h = (0.86, 0.48)

Two perpendicular arrows for the positive and past-tense directions, and the hidden state (0.86, 0.48) as their weighted sum, with dotted lines dropping onto each arrow at 0.9 and 0.4

Reading it: the axes are the two neurons. The blue and green arrows are the two feature directions; the red arrow is the hidden state. The dotted lines drop from the hidden state onto each feature arrow, and where they land (0.9 of the way along blue, 0.4 along green) is the dot product: the amount of that feature. Notice that the red arrow's coordinates on the neuron axes, 0.86 and 0.48, are neither amount. To read a model you have to look along the right directions, not along the neurons.

That features are directions, and that most of what a model tracks can be read with a dot product, is called the linear representation hypothesis. It is a hypothesis, not a law, but it holds often enough to power everything below. You have met it before: word-vector arithmetic such as king − man + woman ≈ queen (primer.ml.embeddings.word2vec) works because "royalty" and "gender" are directions.

In code: compose_features builds a hidden state from feature amounts and read_features reads them back with dot products.

Chapter 2

Probes: can a straight line read it out?

Everyday picture A doctor can't ask your liver how it is doing, but a blood test can measure a marker that tells them. A probe is a blood test for a hidden state: a tiny classifier, trained by you, that looks at a layer's activations and answers one yes/no question, such as "is this review positive?". The model itself is frozen; only the probe learns.

A tiny worked example. A probe for "positive" on a two-neuron layer has weights w = (1, 0.5) and bias b = 0. On the hidden state h = (2, −1):

  1. Score it: 1 · 2 + 0.5 · (−1) + 0 = 1.5.
  2. Squash the score into a probability with the sigmoid σ(z) = 1 / (1 + e^−z), which maps any number into the range 0 to 1: σ(1.5) = 1 / (1 + 0.223) = 0.818.

The probe is 82% sure this hidden state belongs to a positive review. On h = (−1, 1) the score is −1 + 0.5 = −0.5 and σ(−0.5) = 0.378: probably negative.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the frozen model's hidden state for one example (2, −1)
the probe's weights: one per neuron, learned (1, 0.5)
the probe's bias: one number, learned 0
the probe's raw score (its logit) 1.5
the sigmoid: turns any score into a probability between 0 and 1
Euler's number, about 2.718
the probe's probability that the property is present 0.818

In words: "dot the hidden state with the probe's weights, add the bias, and squash the result into a probability."

With the numbers: $p = \sigma(1 \cdot 2 + 0.5 \cdot (-1) + 0) = \sigma(1.5) = 1 / (1 + 0.223) = 0.818$.

Level 3: in Python
import math
w = [1.0, 0.5]
b = 0.0
h = [2.0, -1.0]
# w · h + b
z = sum(wi * hi for wi, hi in zip(w, h)) + b
z  # → 1.5
# σ(z)
round(1 / (1 + math.exp(-z)), 3)  # → 0.818
# a second hidden state, (−1, 1), reads as probably negative
round(1 / (1 + math.exp(-(-1 * 1.0 + 1 * 0.5))), 3)  # → 0.378

This is exactly logistic regression (primer.ml.neural_net), trained the usual way: show it examples whose answer you know, measure its binary cross-entropy (primer.ml.losses), and nudge w and b downhill. The gradient of that loss with respect to the score is simply (p − y), the gap between the probe's probability and the true 0/1 label, which makes each update one line of code.

Figure 3 · Diagram

Reading it: the solid arrows are one forward pass; the dotted arrow is learning. The gradient stops at the probe: the model under study never changes, so whatever the probe finds was already in the hidden state. That is the point of a probe. It measures the model, not a new model trained on top of it.

On our toy model. PlantedModel stores two known features in a 32-neuron hidden layer: sentiment and tense, each as ±1 along its own random direction, plus random noise of size 0.4 on every neuron. No single neuron holds either feature. A probe trained on just 100 examples reads sentiment correctly on 99% of 1,000 examples it never saw.

Why keep probes simple, and check them against a control. A probe can succeed for the wrong reason. Give a probe random labels, a control task with nothing real to find, and it still fits 73% of its 100 training examples, because 33 adjustable numbers can memorize a lot of noise. On unseen examples it scores 50%, a coin flip. Two habits follow: always score a probe on examples it never saw, and compare it with a control task, so that "the probe works" means "the hidden state encodes this" rather than "the probe is clever". This is also why probes are kept linear: a powerful probe (a deep network) could compute the property from raw ingredients by itself, and then its success would tell you about the probe, not the model.

Decodable is not the same as used

Everyday picture A library holds a book nobody ever borrows. Finding it on the shelf proves the library has it, not that it shaped anything any reader did.

A tiny worked example. Our toy model's output reads sentiment and, by construction, gives tense a weight of exactly zero. A probe still reads tense at 98% on unseen examples. Now intervene: take 200 examples, flip only one feature's label, keep everything else fixed, and measure how much the model's output moves.

Flip Probe accuracy for it Output moves by
sentiment 99% 2.00
tense 98% 0.00

Figure 4 · Drawn from the lesson's code

sentiment tense random labels 0.0 0.2 0.4 0.6 0.8 1.0 1.2 probe accuracy What a probe can read 0.99 0.98 0.49 chance training examples unseen examples flip sentiment flip tense 0.0 0.5 1.0 1.5 2.0 2.5 change in the model's output What the model actually uses 2.00 0.00

Left, probe accuracy: sentiment 0.99 and tense 0.98 on unseen examples, random labels 0.73 on training examples but about 0.5 on unseen ones. Right, flipping sentiment moves the output by 2 and flipping tense moves it by 0

Reading it: the left panel is what probes can read. Grey bars are accuracy on the probe's own training examples, blue bars on unseen ones, and the red dashed line is chance. Sentiment and tense are both clearly readable; random labels look readable on the training set and collapse to chance on unseen data, which is exactly what the control is for. The right panel is what the model uses: the change in its output when one feature is flipped. Tense is as readable as sentiment and has no effect at all. The two panels answer different questions, and only the right one is about cause.

A probe finding information does not mean the model uses it. To find out what the model uses you have to change something and watch the output, which is what the rest of this lesson does.

In code: train_probe fits a Probe by gradient descent on frozen hidden states from PlantedModel; probe_report trains the three probes, and flip_effect performs the intervention.

Chapter 3

The logit lens: reading the model's mind mid-thought

Everyday picture A writer keeps every draft of an article. Draft 1 says the landmark is "in a European capital", draft 2 says "probably Paris", the final says "Paris". Reading the drafts shows when the writer made up their mind. The logit lens reads a language model's drafts.

It works because of how a transformer is built (primer.ml.transformer). Each word has a running vector, the residual stream. Every layer reads it and adds a correction to it, rather than replacing it. At the very end, the output layer (the unembedding) turns the final vector into one score per word in the vocabulary. Because every layer writes into the same stream, you can apply that final step early, after any layer, and ask "what would the model predict if it stopped here?"

A tiny worked example. A vocabulary of three words, Paris, London and Rome, a residual stream of two numbers, and an unembedding that scores Paris = first number, London = second number, Rome = minus the first:

After Residual Scores (Paris, London, Rome) Probabilities Top guess
the embedding (0.2, 0.3) (0.2, 0.3, −0.2) (0.36, 0.40, 0.24) London
layer 1, which adds (0.6, 0) (0.8, 0.3) (0.8, 0.3, −0.8) (0.55, 0.34, 0.11) Paris
layer 2, which adds (1.2, −0.3) (2.0, 0.0) (2.0, 0.0, −2.0) (0.87, 0.12, 0.02) Paris

The first draft is a vague "some capital" that happens to lean London; layer 1 tips it to Paris; layer 2 commits.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
which layer we stop after (0 = straight after the embedding) 2
the residual stream after layer , for one word (2.0, 0.0)
the model's final normalization (primer.ml.transformer); our 2-number toy has none, so here LN(h) = h (2.0, 0.0)
the unembedding: one row per vocabulary word, dotted with the state to give that word's score rows (1, 0), (0, 1), (−1, 0)
the scores (logits), one per word (2, 0, −2)
softmax turns scores into probabilities that add up to 1 (primer.ml.attention) (0.87, 0.12, 0.02)
what the model would predict if it stopped after layer Paris at 0.87

In words: "take the residual stream partway up, push it through the model's own final normalization and output layer, and read off the probabilities."

With the numbers: after layer 2, $W_U h_2 = (1 \cdot 2 + 0 \cdot 0,; 0 \cdot 2 + 1 \cdot 0,; -1 \cdot 2 + 0 \cdot 0) = (2, 0, -2)$; then , , , total 8.52, so P(Paris) = 7.39 / 8.52 = 0.87.

Level 3: in Python
import math
# rows: Paris, London, Rome
W_U = [[1, 0], [0, 1], [-1, 0]]
h = [0.2, 0.3]
writes = [[0.6, 0.0], [1.2, -0.3]]
# a residual layer adds its write to the stream
for w in writes:
    h = [a + b for a, b in zip(h, w)]
[round(x, 2) for x in h]  # → [2.0, 0.0]
# W_U h: one score per word
scores = [sum(u * x for u, x in zip(row, h)) for row in W_U]
[round(s, 2) for s in scores]  # → [2.0, 0.0, -2.0]
# softmax
exps = [math.exp(s) for s in scores]
[round(e / sum(exps), 2) for e in exps]  # → [0.87, 0.12, 0.02]

Figure 5 · Diagram

Reading it: the solid line along the middle is the residual stream: it flows from the embedding to the real output, and each layer adds to it. The dotted taps hang the model's own output layer off the stream after every layer. Nothing is trained; the lens borrows weights the model already has. At the top layer the tap and the real output are the same computation, so they must agree exactly, which is a good test that the lens is wired correctly (on primer.ml.transformer.TinyGPT, they do).

On a model where we know the answer. LandmarkModel is a two-layer, one-head transformer built by hand to complete "Eiffel is in" with "Paris". Layer 1 is an MLP that looks up a fact at every word (landmark in, city out); layer 2 is an attention head that lets the last word, "in", find the landmark and copy its city.

Figure 6 · Diagram

Reading it: read left to right as the two steps of the answer. The fact ("the Eiffel Tower is in Paris") is looked up at the landmark's own position in layer 1. It is only moved to the last position, the one that predicts the next word, in layer 2. We built it this way on purpose, so every tool below can be checked against a known truth.

Figure 7 · Drawn from the lesson's code

embedding layer 1 layer 2 0.0 0.2 0.4 0.6 0.8 1.0 lens probability Worked example: the guess firms up Paris London Rome Eiffel is in after embedding after layer 1 (MLP) after layer 2 (attention) Landmark model: P("Paris") by layer and word 0.33 0.33 0.33 0.91 0.33 0.33 1.00 0.69 0.91 0.0 0.2 0.4 0.6 0.8 1.0

Left, the worked example's probabilities for Paris, London and Rome after the embedding, layer 1 and layer 2. Right, a 3 by 3 grid of the lens's probability of Paris by layer and word: the Eiffel column turns dark after layer 1, the in column only after layer 2

Reading it: the left panel is the worked example: at each layer, three bars for the three words, and the blue Paris bar grows from 0.36 to 0.87. The right panel runs the lens over the landmark model: rows are layers, columns are words, and each cell is the lens's probability of "Paris" (1/3 = 0.33 means no opinion among three cities). The Eiffel column goes dark after layer 1, when the MLP looks up the fact. The "in" column only goes dark after layer 2, when attention copies the city over. The lens has shown where and when the answer appears, without being told.

The logit lens was first described for GPT-2 in 2020, where middle layers already "guess" the next word surprisingly well. In some models the early layers read as nonsense through the final output layer, because they do not yet speak its "language"; the tuned lens fixes this by training a small translator for each layer before applying the output layer.

In code: logit_lens applies an output layer to any residual states, worked_lens runs the table above, residual_stream and tinygpt_logit_lens do the same for primer.ml.transformer.TinyGPT, and lens_map builds the grid for LandmarkModel.

Chapter 4

Activation patching: which activations cause the answer?

Everyday picture Two cars of the same model sit side by side: one starts, one doesn't. You move parts from the good car into the bad one, one part at a time, and try the ignition after each swap. The part whose swap makes the bad car start is the part that mattered. Activation patching (also called causal tracing) does this with a model's internal activations.

A tiny worked example. Run the landmark model twice:

  • the clean prompt "Eiffel is in": scores Paris 3, Rome 0, so the logit difference Paris − Rome is +3;
  • the corrupted prompt "Colosseum is in": Paris 0, Rome 3, so the difference is −3.

Now run the corrupted prompt again, but at one chosen place overwrite the activation with the one from the clean run, and see how much of the gap between −3 and +3 comes back:

  • patch the MLP output at the landmark's position: the difference jumps back to +3, so 100% of the answer is restored;
  • patch the state at "is": it stays at −3, 0% restored ("is" is the same word in both prompts, so its state carries nothing about the landmark).
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
logit difference: the correct answer's score minus the wrong answer's score (Paris − Rome)
on the clean prompt, nothing patched +3
on the corrupted prompt, nothing patched −3
on the corrupted prompt, with one clean activation copied in +3, −3, or anything between
restored the fraction of the gap that one patch closes: 0 = nothing, 1 = everything 1, 0

In words: "how far the patch moved the answer, as a share of the full distance from the corrupted answer to the clean one."

With the numbers: patching the landmark's MLP output gives (3 − (−3)) / (3 − (−3)) = 6 / 6 = 1. Patching "is" gives (−3 − (−3)) / 6 = 0. A patch that left the difference at 0 would give (0 − (−3)) / 6 = 0.5: half the answer restored.

Level 3: in Python
clean, corrupt = 3.0, -3.0
# restored = (LD_patched − LD_corrupt) / (LD_clean − LD_corrupt)
def restored(patched):
    return (patched - corrupt) / (clean - corrupt)
# the landmark's MLP output, patched in
restored(3.0)  # → 1.0
# the state at "is", patched in
restored(-3.0)  # → 0.0
# a patch that only reaches a tie
restored(0.0)  # → 0.5

Why a difference of two scores, rather than the probability of Paris? Because the corrupted prompt differs from the clean one in exactly one fact, the difference measures exactly that fact, and it moves smoothly, where a probability can saturate near 0 or 1 and hide a change.

Figure 8 · Diagram

Reading it: there are three runs. The clean run is a donor: its activations are saved. The corrupted run sets the baseline. The patched run is the corrupted prompt with a single activation transplanted from the donor. Repeat the patched run once per layer and position and you get one "fraction restored" per site: a map of where the answer is carried.

Figure 9 · Drawn from the lesson's code

Eiffel / Colosseum is in position patched (clean / corrupted word) after embedding after layer 1 (MLP) after layer 2 (attention) Fraction of the answer restored by one patch 1.00 0.00 0.00 1.00 0.00 0.00 0.00 0.00 1.00 0.0 0.2 0.4 0.6 0.8 1.0

A 3 by 3 grid, layers by positions, of the fraction of the answer restored by one patch: 1 at the Eiffel column for the embedding and layer 1, 1 at the in column after layer 2, 0 everywhere else

Reading it: rows are the residual stream after each layer, columns are the three positions (clean word / corrupted word), and each cell is the fraction of the answer restored by patching that one state. The answer lives at the landmark early (the Eiffel column is 1 after the embedding and after the MLP) and at the last word late (the "in" column is 1 after attention). The diagonal hand-off between them is the two-step circuit we built, found by intervention alone.

Compare with the lens grid above. After layer 2, the lens reads "Paris" at the Eiffel position with probability 0.995, yet patching that state restores 0.00: after the last layer nothing reads that position any more. The information is there, and it no longer matters. Readable is not the same as used, again.

Why it matters in practice. This is how researchers located where GPT-style models store facts: causal tracing on prompts like ours found factual recall concentrated in middle-layer MLPs at the subject's last token, and then used that to edit single facts. The same method, one site at a time, has traced whole circuits, such as the one GPT-2 small uses to fill in "When Mary and John went to the store, John gave a drink to" → "Mary".

In code: LandmarkModel is the hand-built model and LandmarkModel.run accepts patches; fraction_restored is the formula and patching_map patches every layer and position in turn.

Chapter 5

Superposition: more features than neurons

Everyday picture Back to the band, but now five instruments play through two speakers. The engineer pans each instrument to its own position around the room: hard left, front right, back left, and so on. When one instrument plays alone, you can tell which one by where the sound comes from. When two play at once, the positions blur together and you might mistake the pair for a third instrument. The trick works because in this band, most of the time, only one instrument is playing.

Models face the same squeeze. There are far more concepts in the world than neurons in a layer, but in any one sentence almost all of them are absent: features are sparse. So models store more features than they have dimensions, as directions that are nearly perpendicular rather than exactly. That is superposition.

A tiny worked example. Put five features in two neurons, spread 72° apart like the points of a pentagon. Feature 's direction is (cos 72k°, sin 72k°): feature 0 is (1, 0), feature 1 is (0.309, 0.951), and so on. Switch on feature 0 alone, at strength 1. The hidden state is (1, 0). Reading every feature back with a dot product gives

Feature Angle from feature 0 Read-back (cos of the angle) After adding −0.31 and ReLU
0 0° 1.000 0.69
1 72° 0.309 0
2 144° −0.809 0
3 216° −0.809 0
4 288° 0.309 0

Five directions can't all be perpendicular in two dimensions, so reading feature 0 leaks 0.309 into each neighbour: interference. The fix is a small negative bias and a ReLU (which turns negatives into zero): the leak of 0.309 falls below the bias of 0.31 and is filtered to 0, and the real feature survives at 0.69. The cost comes when two neighbours are on at once: each then reads 1 + 0.309 − 0.31 = 0.999 instead of 0.69, too high. Sparsity is a bet that such collisions are rare.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the true features: one strength per feature (1, 0, 0, 0, 0)
the squeeze: 2 rows (neurons) by 5 columns (features); column is feature 's direction the pentagon
the hidden state: 2 numbers holding 5 features (1, 0)
transposed (rows and columns swapped): reads each feature back with a dot product
one bias per feature, learned; negative values filter small leaks −0.31 each
ReLU keep positives, turn negatives into 0
the features read back out (0.69, 0, 0, 0, 0)
feature 's direction dotted with itself (its squared length) 1
how much feature leaks into feature : 0 if perpendicular 0.309 for neighbours
add up over every other feature

In words: "squeeze the features into a few neurons, read each one back by dotting with its direction, subtract a threshold and drop anything negative. What you read for feature is its own strength, plus a leak from every other active feature, minus the threshold."

With the numbers: with feature 0 alone on, $\hat{x}_1 = \text{ReLU}(1 \cdot 0 + 0.309 \cdot 1 - 0.31) = \text{ReLU}(-0.001) = 0$ and . With features 0 and 1 both on, .

Level 3: in Python
import math
# feature k's direction: (cos 72k°, sin 72k°)
W = [[math.cos(math.radians(72 * k)), math.sin(math.radians(72 * k))] for k in range(5)]
def read_back(x, bias, relu=True):
    # h = W x: two numbers
    h = [sum(W[k][n] * x[k] for k in range(5)) for n in range(2)]
    # Wᵀ h + b: one number per feature
    z = [W[k][0] * h[0] + W[k][1] * h[1] + bias for k in range(5)]
    # ReLU keeps positives and turns negatives into 0
    return [round(max(0.0, v) if relu else v, 3) for v in z]
# the leak: feature 0 alone, no bias, no ReLU
read_back([1, 0, 0, 0, 0], bias=0, relu=False)  # → [1.0, 0.309, -0.809, -0.809, 0.309]
# the bias and ReLU filter it
read_back([1, 0, 0, 0, 0], bias=-0.31)  # → [0.69, 0.0, 0.0, 0.0, 0.0]
# two neighbours on: each reads too high
read_back([1, 1, 0, 0, 0], bias=-0.31)  # → [0.999, 0.999, 0.0, 0.0, 0.0]

Figure 10 · Diagram

Reading it: follow a feature through the bottleneck. Five numbers go in, are squeezed into two, and must be expanded back into five. There is no way to fit five perpendicular directions into two neurons, so the read-back always leaks; the bias and ReLU at the end are what make the leak survivable, by throwing away small readings. This is the toy model from Anthropic's Toy Models of Superposition (2022), and the next step is to train it and see what it chooses.

Training it. Let the model learn and itself, by gradient descent on the reconstruction error, with features that matter less and less (feature is weighted ). Compare two worlds:

  • Dense: every feature is on in every example. The model keeps the two most important features, perpendicular, at length 1.00, and gives the other three length 0: it simply drops them.
  • Sparse: each feature is on only 5% of the time. The model keeps all five, 72° apart, at length about 1.1, and learns a negative bias of about −0.23 to filter the leaks.

Figure 11 · Drawn from the lesson's code

−1.0 −0.5 0.0 0.5 1.0 neuron 1 −1.0 −0.5 0.0 0.5 1.0 neuron 2 Dense features: keeps 2 (each feature on 100% of the time) f0 f1 −1.0 −0.5 0.0 0.5 1.0 neuron 1 −1.0 −0.5 0.0 0.5 1.0 neuron 2 Sparse features: keeps all 5 (each feature on 5% of the time) f0 f1 f2 f3 f4 f0 f1 f2 f3 f4 −1.00 −0.75 −0.50 −0.25 0.00 0.25 0.50 0.75 1.00 how much the feature moves the neuron Each neuron answers to several features neuron 1 neuron 2

Three panels. Left, dense features: two perpendicular arrows and three near-zero ones. Middle, sparse features: five arrows spread evenly around a pentagon. Right, bars of how much each of five features moves neuron 1 and neuron 2

Reading it: in the two left panels each arrow is one feature's learned direction in the two-neuron space, and its length is how well the model stores it. Dense features (left) get the textbook answer: as many features as neurons, perpendicular, the rest discarded. Sparse features (middle) get a pentagon: five features in two dimensions, accepting a little interference in exchange for storing everything. The right panel reads the ideal pentagon by neuron: neuron 1 moves for feature 0 (by 1.0), feature 1 and feature 4 (by 0.309 each), and moves the other way for features 2 and 3. No neuron belongs to one feature.

That last point is why looking at single neurons so often fails. A neuron that responds to several unrelated features is polysemantic. Vision researchers found neurons that respond to cat faces and to the fronts of cars; language models are full of neurons like that. Superposition explains why: when features outnumber neurons, the features cannot line up with the neurons, so every neuron is a mixture.

In code: pentagon builds the five directions, superposition_readout is the formula, train_superposition learns W and b with gradients from superposition_loss_and_grads (checked against finite_difference_gradient), and features_a_neuron_responds_to lists a neuron's features.

Chapter 6

Sparse autoencoders: getting the features back

Everyday picture A sound engineer receives the two-speaker recording of the five-instrument band, with no notes on who played when. She knows one thing about this band: usually only one instrument plays at a time. Many different scores could produce the same sound, but she writes down the one that uses the fewest instruments. That preference is what lets her recover the real parts rather than some arbitrary mixture.

A sparse autoencoder (SAE) does this for a model's hidden states. It learns a dictionary: many more directions than the layer has neurons (it is overcomplete), trained so that every hidden state can be rebuilt from just a few of them. Each learned direction is a candidate feature, and each one's strength on an input is its latent activation.

A tiny worked example. The hidden state is h = (0.8, 0.6) and the dictionary has three directions, (1, 0), (0, 1) and (0.8, 0.6). Two codes rebuild h perfectly:

Code (strength of each direction) Rebuilt Error Penalty with λ = 0.1 Loss
A: (0.8, 0.6, 0), two directions (0.8, 0.6) 0 0.1 × (0.8 + 0.6) = 0.14 0.14
B: (0, 0, 1), one direction (0.8, 0.6) 0 0.1 × 1.0 = 0.10 0.10

Rebuilding alone can't choose between them. The penalty on the total strength, the L1 penalty (primer.ml.regularization), prefers the code that uses one direction. One side effect: the best strength for direction 3 is not quite 1. With strength , the loss is , which is lowest at (loss 0.0975): the penalty always pulls strengths a little toward zero, known as shrinkage.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
a hidden state from the model under study (0.8, 0.6)
how many latents (dictionary directions) the SAE has, usually far more than neurons 3
, the encoder's weights and biases: turn a hidden state into latent strengths
the code: one strength per latent, ReLU keeps them ≥ 0 and mostly exactly 0 (0, 0, 1)
the decoder: column is latent 's direction, kept at length 1 (1, 0), (0, 1), (0.8, 0.6)
the decoder bias: the typical hidden state, subtracted before encoding and added back after (0, 0)
the rebuilt hidden state (0.8, 0.6)
squared length of the rebuild error: add up the squares of its entries 0
the sparsity penalty's strength, chosen by you 0.1
the size of latent 's strength
the loss to minimize: rebuild error plus penalty 0.10

In words: "encode the hidden state into many non-negative strengths, decode it back as a weighted sum of dictionary directions, and pay for both the rebuild error and the total strength used."

With the numbers: for code B, $\hat{h} = 1 \cdot (0.8, 0.6) = (0.8, 0.6)0.1 \times 1 = 0.10L = 0.10$. For code A the penalty is .

Level 3: in Python
dictionary = [[1, 0], [0, 1], [0.8, 0.6]]
h = [0.8, 0.6]
lam = 0.1
def loss(code):
    # ĥ = Σ f_i · direction_i
    rebuilt = [sum(f * d[n] for f, d in zip(code, dictionary)) for n in range(2)]
    # ‖h − ĥ‖² + λ Σ |f_i|
    return round(sum((a - b) ** 2 for a, b in zip(h, rebuilt)) + lam * sum(abs(f) for f in code), 4)
loss([0.8, 0.6, 0])  # → 0.14
loss([0, 0, 1.0])  # → 0.1
# shrinkage: a slightly weaker code is cheaper still
loss([0, 0, 0.95])  # → 0.0975

Figure 12 · Diagram

Reading it: the hidden state is blown up into a much wider code and squeezed back down. The loss watches two things at once: the rebuild must match the original (the arrow from h), and the code must be small (the arrow from f). Without the second arrow the SAE could use any directions at all. With it, each input is explained by a handful of latents, and each latent's decoder column is a candidate feature you can inspect: read which inputs make it fire, and add its direction to the model to see what it does.

On the superposition toy. Take 20,000 hidden states from the pentagon model (each feature on 5% of the time), train an SAE with 5 latents and λ = 0.3, and compare the learned decoder columns with the five planted directions. Every planted feature is matched by a latent with cosine similarity above 0.99 (1 would be the same direction), and an example with a feature on lights up about 1.07 latents on average. The price is shrinkage: 11% of the variance goes unexplained. With a tiny penalty, λ = 0.01, the SAE rebuilds the data perfectly (0.0% unexplained), yet its worst match to a real feature is only 0.80: any two directions can rebuild a two-dimensional space, so without sparsity pressure nothing forces the dictionary to find the model's actual features.

Figure 13 · Drawn from the lesson's code

−1.0 −0.5 0.0 0.5 1.0 −1.0 −0.5 0.0 0.5 1.0 λ = 0.3: each latent lands on a feature hidden states planted features learned latents 1 0 − 2 1 0 − 1 1 0 0 sparsity penalty λ 0.0 0.2 0.4 0.6 0.8 1.0 The penalty trades rebuild for meaning worst feature match (cosine) variance left unexplained

Left, grey dots of hidden states lying mostly along five rays, with thick grey planted directions and red learned latent arrows lying on top of them. Right, as the penalty grows from 0.003 to 1, the worst feature match peaks near 0.99 at a penalty of 0.3, while the unexplained variance climbs from 0 to above 0.5

Reading it: on the left, each grey dot is one hidden state. Because features are sparse, most dots lie along one of five rays: one feature on, at some strength. The few dots off the rays are the rare examples with two features on at once. The thick grey lines are the directions we planted; the red arrows are what the SAE learned from the dots alone, and they land on the rays. On the right, the x-axis is the penalty λ on a log scale. The blue line is the worst match between a planted feature and its nearest latent: it sits around 0.8 to 0.9 for small penalties and peaks near 0.99 at λ = 0.3. The red line, the variance left unexplained, is near 0 for small penalties and climbs past 0.5 at λ = 1, where the penalty starts crushing the codes and the match falls again. Choosing λ is a trade: too small and the latents are not the features, too large and too much of the model's activity is lost.

Why it matters in practice. This is the tool that turned interpretability from "a few hand-picked neurons" into something that scales. Anthropic's Towards Monosemanticity (2023) trained SAEs on a small transformer and found thousands of features that each fire for one recognizable thing. Scaling Monosemanticity (2024) did the same inside a production model, Claude 3 Sonnet, and found features for concepts such as the Golden Gate Bridge. Turning that one feature up made the model bring up the bridge in almost every answer, which is the causal check that the feature means what it seems to mean.

In code: sae_objective computes the loss for one proposed code, SparseAutoencoder holds the encoder and decoder with its hand-derived SparseAutoencoder.loss_and_grads, train_sae fits it, and match_features compares learned directions with planted ones.

Chapter 7

What these tools cannot (yet) show

Everyday picture A brain scan shows which regions light up while a person reads, but not the sentence they are thinking. Every tool in this lesson is like that: a real instrument that measures something true and partial.

A tiny worked example. Each of our own toys already showed a limit:

What we saw The limit it shows
A probe read tense at 98%, and the output ignored tense completely Probes find information, not use
The lens read "Paris" at a position where patching restored 0% Reading an activation doesn't mean anything downstream reads it
The weak SAE rebuilt everything perfectly and still found the wrong directions A good rebuild score doesn't mean the features are real
The strong SAE left 11% of the variance unexplained Some of the model's activity is in no feature the dictionary found
We knew the landmark model's circuit because we built it In a real model, nobody has the answer key

Figure 14 · Diagram

Reading it: the arrow is the direction of stronger evidence. Most claims about real models sit on the left or in the middle. A complete mechanism, every step from input to output accounted for, exists only for narrow behaviours in small or carefully chosen settings. When you read an interpretability result, ask which box it reached.

The main open problems, in plain words:

  • Choices shape answers. Patching results depend on how the corrupted prompt is built (swap one word? add noise?). SAE features depend on the dictionary's size and λ: a bigger dictionary can split one feature into several finer ones.
  • Redundancy hides importance. Patching one site at a time can miss parts that back each other up. Models have been observed to partly repair themselves when one component is knocked out, so "patching this restores 0%" does not always mean "this plays no role".
  • Coverage. The unexplained part of an SAE's rebuild is not noise to ignore: it is model behaviour no feature describes yet.
  • Scale and labour. A circuit that explains one behaviour of a small model can take researchers weeks. No one has a complete account of how a frontier model produces any long answer.
  • Labels are human guesses. Naming a feature "Golden Gate Bridge" summarizes the inputs that make it fire; checking that the name is right needs interventions like the one above.

Why it matters in practice. Treat these tools as evidence, stacked next to behavioural evaluations (primer.agents.evals), never as a certificate. A probe or a lens is a cheap first look; a patching or steering experiment is the minimum for a causal claim; and a clean result on a toy, like every result in this lesson, is where understanding starts, not where it ends.

Test yourself

6 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1What does it mean to say a feature is a "direction" rather than a neuron?Think it through, then reveal

The feature's presence is stored as a pattern across many neurons: the hidden state moves along a particular direction in proportion to how strongly the feature is present. You read it with a dot product against that direction. Any single neuron is typically a mix of several features.

Question 2A linear probe reads a property from layer 12 at 95% accuracy. What can you conclude, and what can't you?Think it through, then reveal

You can conclude the property is linearly decodable from layer 12, as long as the 95% was on held-out examples and clearly beats a control task with random labels. You can't conclude the model uses it. For that you need an intervention: remove or change that information and see whether the output changes.

Question 3How does the logit lens work, and why does it agree with the model at the last layer?Think it through, then reveal

It takes the residual stream after some layer and applies the model's own final normalization and unembedding, turning it into next-token probabilities. At the last layer that is exactly the computation the model itself performs, so the two must match. Earlier layers are read "as if the model stopped there".

Question 4Describe an activation patching experiment, and what the "fraction restored" means.Think it through, then reveal

Run a clean prompt and save its activations; run a corrupted prompt that changes one fact; then rerun the corrupted prompt with one activation replaced by its clean value. The fraction restored is how much of the clean-minus-corrupted logit difference the patch brings back: 1 means that activation alone carries the fact, 0 means it carries none of it.

Question 5Why do models use superposition, and why does it make neurons hard to interpret?Think it through, then reveal

There are more useful features than neurons. When features are rarely active at the same time, the model can store them as nearly perpendicular directions and filter the small interference with a bias and ReLU. The directions can't line up with the neurons, so each neuron responds to several unrelated features: it is polysemantic.

Question 6Why does a sparse autoencoder need the L1 penalty? What goes wrong if λ is too small or too large?Think it through, then reveal

Many dictionaries rebuild the data equally well; the penalty picks the one where each input uses few latents, which pushes latents onto the real features. Too small and the SAE rebuilds perfectly with meaningless directions; too large and it shrinks activations and leaves much of the model's activity unexplained.

Primary sources

The papers behind this lesson

Alain & Bengio, Understanding intermediate layers using linear classifier probes (2016)

Introduced linear probes as a way to measure what each layer of a network makes linearly available.

The paper ↗
Hewitt & Liang, Designing and Interpreting Probes with Control Tasks (2019)

Showed that probes can succeed by memorizing, and introduced control tasks and selectivity to tell the two apart.

The paper ↗
Belrose et al., Eliciting Latent Predictions from Transformers with the Tuned Lens (2023)

Formalized the logit lens and fixed its failures on early layers by training a small translator for each layer.

The paper ↗
Meng et al., Locating and Editing Factual Associations in GPT (2022)

Introduced causal tracing, found factual recall in middle-layer MLPs at the subject's last token, and edited single facts there.

Read the annotated companion →The paper ↗
Wang et al., Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small (2022)

Used patching to reverse-engineer a complete circuit for one behaviour of a real language model.

The paper ↗
Elhage et al., Toy Models of Superposition (2022)

Showed with small ReLU models that sparse features are stored in superposition, and when and how the geometry changes.

Read the annotated companion →The paper ↗
Bricken et al., Towards Monosemanticity: Decomposing Language Models With Dictionary Learning (2023)

Trained sparse autoencoders on a small transformer and found thousands of interpretable features hidden in polysemantic neurons.

Read the annotated companion →The paper ↗
Templeton et al., Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet (2024)

Scaled sparse autoencoders to a production model and steered its behaviour through the features they found.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Olah et al., Zoom In: An Introduction to Circuits (Distill, 2020): https://distill.pub/2020/circuits/zoom-in/
  • Elhage et al., A Mathematical Framework for Transformer Circuits (2021): https://transformer-circuits.pub/2021/framework/index.html
  • Elhage et al., Toy Models of Superposition (2022): https://transformer-circuits.pub/2022/toy_model/index.html
  • Bricken et al., Towards Monosemanticity (2023): https://transformer-circuits.pub/2023/monosemantic-features/index.html
  • Templeton et al., Scaling Monosemanticity (2024): https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html
  • Cunningham et al., Sparse Autoencoders Find Highly Interpretable Features in Language Models (2023): https://arxiv.org/abs/2309.08600
  • Belinkov, Probing Classifiers: Promises, Shortcomings, and Advances (2021): https://arxiv.org/abs/2102.12452
  • Meng et al., Locating and Editing Factual Associations in GPT (2022): https://arxiv.org/abs/2202.05262
  • Belrose et al., Eliciting Latent Predictions from Transformers with the Tuned Lens (2023): https://arxiv.org/abs/2303.08112

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.