At a glance
Key takeaways
- The old way read words in order, passing a running summary from each word to the next. Slow to train, and forgetful over long distances.
- The new way lets every word ask a question of every other word and blend in the answers, all at once. That asking and blending is attention.
- Stack it: attention, then a little private thinking per word, repeated six times, with shortcuts around every step so training stays stable.
- Stamp positions on: attention ignores order, so each word gets a pattern of waves that encodes where it sits.
- The payoff: better translations than any earlier system, for a fraction of the training compute. Within three years the design had spread to images, audio and code.
Chapter 1 Paper §1–2 ↗
Reading through a keyhole
Picture a relay race where each runner can start only when the one before hands over the baton. Adding runners doesn't make the race any faster. That is how a recurrent neural network reads a sentence: to process word 50 it must first finish words 1 to 49, because each step needs the summary produced by the step before.
Two problems follow. A GPU, built to do thousands of things at once, sits mostly idle while it waits. And what the first word said reaches the last word only after surviving every hand-off in between, so long-range connections fade.
Figure 1 · Interactive
The relay race and the round table
Others had tried to escape the relay. Convolutional models read words through small overlapping windows, which run in parallel, but two distant words only meet after many stacked layers. The Transformer's bet was more radical: drop the step-by-step memory entirely, and let every word connect to every other word directly.
“This inherently sequential nature precludes parallelization within training examples.”Vaswani et al. (2017), §1
Chapter 2 Paper §3.2 ↗
Every word looks at every word
You walk into a library and type a search: your query. Every book has a catalogue card describing it, its key, and the book itself on the shelf, its value. The librarian compares your search with every card and hands you not one book but a blend: mostly the best match, a little of the next best.
In a Transformer, every word plays all three parts at once. Each word searches, is catalogued and gets borrowed. This is self-attention, and it's the only place in the whole model where words exchange information.
“An attention function can be described as mapping a query and a set of key-value pairs to an output.”Vaswani et al. (2017), §3.2
Chapter 3 Paper §3.2.1 ↗
One attention step, by hand
Let's do the whole calculation for the word it, with numbers small enough to check in your head. Each word gets three short lists of numbers (vectors): a query, a key and a value. Real models learn these, and use 64 numbers each. We'll use four, and two for the values so we can draw them.
The formula, decoded
In words: score every query against every key, shrink the scores by the square root of the key length, turn each row into shares, and use the shares to blend the values.
The same numbers in Python
import math
q = [1, 1, 1, 1] # the query for "it"
keys = {"animal": [1, 1, 1, 1], "tired": [1, 1, 0, 0], "street": [1, 0, 0, 0]}
values = {"animal": [1, 0], "tired": [0, 1], "street": [1, 1]}
scores = {w: sum(a * b for a, b in zip(q, k)) / math.sqrt(len(q)) for w, k in keys.items()}
total = sum(math.exp(s) for s in scores.values())
shares = {w: math.exp(s) / total for w, s in scores.items()}
# {'animal': 0.63, 'tired': 0.23, 'street': 0.14}
out = [sum(shares[w] * values[w][i] for w in values) for i in range(2)]
# [0.77, 0.37]
Footnote: why divide by √dk?
Roll one die and the result swings between 1 and 6. Add up a hundred dice and the total swings by dozens. A dot product is a sum of dk small products, so the longer the vectors, the wilder the scores. Wild scores make softmax hand nearly everything to one word, and then it stops responding to small changes. The gradient that training depends on vanishes.
Figure 3 · Interactive · computed live
Softmax with and without the scaling
“We suspect that for large values of dk, the dot products grow large in magnitude, pushing the softmax function into regions where it has extremely small gradients.”Vaswani et al. (2017), §3.2.1
Chapter 4 Paper §3.2.2 ↗
Many heads, many questions
One attention can only ask one kind of question at a time. So give the same sentence to eight readers, each with a different coloured highlighter. One marks who did what, another which pronoun points where, another the word just before. Afterwards, staple their notes together. That's multi-head attention: eight small attentions running side by side, each free to learn its own question.
Figure 4 · Gallery
Six heads, six habits
Rows and columns run through the same eight words: The animal didn't cross because it was tired. Hover a cell to read it.
Splitting 512 into 8 × 64
The paper's word vectors hold 512 numbers. Rather than one attention over all 512, it runs 8 heads over 64 numbers each, glues the eight results back together, and mixes them with one more matrix. The total work is about the same as a single full-width head.
In words: for each head, project the queries, keys and values down to 64 numbers with that head's own matrices and run attention; then place the eight outputs side by side and mix them with one final matrix WO.
Counting weights: each of WQ, WK, WV is 8 heads × (512 × 64), and WO is 512 × 512, so one attention layer holds 4 × 512 × 512 = 1,048,576 projection weights.
The paper's Table 3 tested the trade-off. With the total width fixed, a single head scored 0.9 BLEU worse than eight, and 32 narrow heads also did worse: more heads help, but only up to a point. Chapter 10 charts it.
“Multi-head attention allows the model to jointly attend to information from different representation subspaces at different positions.”Vaswani et al. (2017), §3.2.2
Chapter 5 Paper §3.2.3 ↗
No peeking
The translation model has two halves. The encoder reads the whole English sentence and writes rich notes about every word. The decoder writes the German sentence one word at a time. The same look-around-and-blend tool is used in three places, differing only in who asks and who answers:
1 · Encoder self-attention
The reader rereads their own notes. Every English word sees every other English word.
2 · Masked decoder self-attention
The writer rereads their draft, but never peeks at words not yet written.
3 · Encoder–decoder attention
The writer consults the reader's notes: German queries, English keys and values.
During training the model sees the whole correct German sentence at once, which is what makes training parallel. To stop the decoder cheating by reading the word it's meant to predict, the future is blanked out with a causal mask.
Figure 5 · Interactive
The causal mask
Hover or tap a cell.
Chapter 6 Paper §3.5 ↗
Where am I?
Attention has a blind spot: it has no sense of order. “Dog bites man” and “man bites dog” contain the same words, so without help they'd get the same result. The fix is to stamp each word with its position before anything else happens.
The paper's stamp is a wall of clocks whose hands turn at different speeds: a hand that spins every few words, one that takes dozens, one that takes thousands. Read all the hands together and you know exactly where you are. Better still, moving three words along always turns each hand by the same amount, which makes “three words later” easy for the model to recognise.
Figure 6 · Interactive · computed live
The wall of clocks
dimension 0 (fast) →dimension 127 (slow)
In words: for position pos, each pair of numbers gets the sine and cosine of the position divided by a number that grows from 1 to 10,000 across the pairs. Early pairs spin fast; late pairs spin slowly. The stamp is added to the word's embedding.
The paper also tried learned position vectors and got nearly identical results (Table 3, row E: 25.7 BLEU against 25.8). Most modern models use RoPE instead, which rotates queries and keys by an angle set by position: the same clock-hand idea, applied inside attention.
Chapter 7 Paper §3 ↗
The whole machine
Now assemble the parts. Words become vectors, positions are stamped on, and then the same layer repeats six times: attention, where words talk to each other, then a small feed-forward network, where each word thinks privately about what it heard. Around every step runs a shortcut, the residual connection: each layer writes corrections in the margin rather than rewriting the page. Then layer normalization tidies the numbers so nothing grows out of control.
Figure 7 · The paper's Figure 1, redrawn · Interactive
Follow a sentence through the Transformer
Scroll the diagram sideways to see the decoder →
Start here
Press Follow the data
Or pick any block. We'll translate “The cat sat.” into German with the paper's base model: vectors of 512 numbers, 8 heads, 6 layers.
The two pieces that aren't attention
Feed-forward: widen each word's 512 numbers to 2,048, zero out the negatives (ReLU), and narrow back to 512. The same recipe runs on every word separately.
Add & norm: add the step's proposed change to its input, then rescale to zero mean and unit spread. With x = (2, 0) and a proposed change of (1, 1): (3, 1), which normalizes to (1, −1).
Per layer, the feed-forward network holds 2 × 512 × 2,048 + 2,048 + 512 = 2,099,712 parameters, about twice the attention projections. Across a Transformer, feed-forward layers hold roughly two thirds of the weights. The paper also shares one embedding matrix between the input, the output and the final linear layer, multiplying the embeddings by √512 ≈ 22.6 on the way in.
Chapter 8 Paper §4 ↗
Why it was faster
How many phone calls does it take for news to travel between two people? Along a chain of friends, as many calls as there are people between them. On a group call, one. The paper compares layer types on exactly this, plus how much work each does.
| Layer | Work per layer | Steps in sequence | Hops between two words |
|---|---|---|---|
| Self-attention | O(n²·d) | O(1) | O(1) |
| Recurrent | O(n·d²) | O(n) | O(n) |
| Convolutional | O(k·n·d²) | O(1) | O(logk n) |
Figure 8 · Interactive · computed live
When is n² cheaper than d²?
Chapter 9 Paper §5 ↗
Training it
A new driver pulls out of the driveway slowly, speeds up on the open road, then eases off near the destination. The learning rate follows the same shape. It starts tiny because an untrained model's gradients point in wild directions, climbs during a warmup, then decays so the model can settle into fine detail. The optimizer is Adam.
Figure 9 · Interactive · computed live
The warmup schedule
Hover the curve to read it.
In words: for the first warmup steps, grow the learning rate in a straight line; after that, shrink it as one over the square root of the step number; scale everything by one over the square root of the model width.
Two more tricks fight overfitting. Dropout (P = 0.1) zeroes a random tenth of each sub-layer's output during training. Label smoothing (ε = 0.1) marks the right answer as 90% right rather than 100%. The paper notes that label smoothing makes perplexity worse, because the model learns to be less sure, but improves BLEU: a training metric and a quality metric disagreeing.
Chapter 10 Paper §6 ↗
Better, and cheaper
BLEU scores a translation by how many of its word sequences also appear in a professional translation, on a 0 to 100 scale. A gain of one or two points counts as a clear step forward. The Transformer didn't just win; it won while costing less to train.
Figure 10 · Table 2, charted
Quality against training cost, English→German
Figure 11 · Table 3 (A), charted
How many heads?
Chapter 11 Paper §7 ↗
What happened next
The paper closes by planning to take attention to images, audio and video, to study attention restricted to local windows for long inputs, and to make generation less sequential. All three happened. The core (attention and feed-forward layers wrapped in residual connections) is unchanged to this day. Nearly every detail around it has been tuned.
What changed since 2017
| In the paper | Common today | Why |
|---|---|---|
| Encoder + decoder | Decoder-only for chat; encoder-only for embeddings | One stack of masked attention generates text; a two-way stack makes the best embeddings |
| Norm after each step | Norm before each step, often RMSNorm | More stable training at depth |
| Sine and cosine positions | RoPE | Relative positions, and easier to stretch to longer contexts |
| ReLU feed-forward | GELU or SwiGLU, sometimes mixture of experts | Better quality per parameter |
| 8 heads, each with its own keys and values | Grouped-query attention | A smaller KV cache, so more users per GPU |
| Plain attention | FlashAttention kernels | The same math with far less memory traffic |
| 6 layers, 65M to 213M parameters | Dozens of layers, billions of parameters | Scaling laws: bigger models on more data keep improving |
“In this work, we presented the Transformer, the first sequence transduction model based entirely on attention.”Vaswani et al. (2017), §7
Test yourself
Five questions
Each one checks an idea, not a fact you could look up. Pick an answer to see why it's right or wrong.
Build it yourself
Every piece, in plain Python.
The AI Primer builds each idea on this page from scratch with NumPy alone, with tests that prove it works. Its lessons are being illustrated here, starting with attention.
About this page. This article explains the paper; it doesn't reproduce it. It quotes a sentence at most per chapter, clearly marked, and every figure is redrawn from scratch. The tables are reproduced in part with attribution under the notice printed on the paper: “Provided proper attribution is provided, Google hereby grants permission to reproduce the tables and figures in this paper solely for use in journalistic or scholarly works.” Numbers marked illustrative were chosen to teach; everything else comes from the paper or is computed live from its formulas. Read the original alongside: every chapter links to its section.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł. and Polosukhin, I. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems 30. arXiv:1706.03762.