rumblr Work in progressWIP

● Papers · 2017 · NeurIPS

Attention Is All You Need

Before 2017, language models read a sentence one word at a time, like reading through a keyhole. This paper let every word look at every other word at once. The design it called the Transformer is the blueprint inside GPT, Claude, Gemini and nearly every model that followed.

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin
Begin The paper on arXiv
How it works tells the story and decodes every formula. Math & code adds the paper’s fine print.

Above: which words the word it attends to. The pattern is illustrative, set by hand.

At a glance

Key takeaways

  1. The old way read words in order, passing a running summary from each word to the next. Slow to train, and forgetful over long distances.
  2. The new way lets every word ask a question of every other word and blend in the answers, all at once. That asking and blending is attention.
  3. Stack it: attention, then a little private thinking per word, repeated six times, with shortcuts around every step so training stays stable.
  4. Stamp positions on: attention ignores order, so each word gets a pattern of waves that encodes where it sits.
  5. The payoff: better translations than any earlier system, for a fraction of the training compute. Within three years the design had spread to images, audio and code.

Chapter 1 Paper §1–2 ↗

Reading through a keyhole

Picture a relay race where each runner can start only when the one before hands over the baton. Adding runners doesn't make the race any faster. That is how a recurrent neural network reads a sentence: to process word 50 it must first finish words 1 to 49, because each step needs the summary produced by the step before.

Two problems follow. A GPU, built to do thousands of things at once, sits mostly idle while it waits. And what the first word said reaches the last word only after surviving every hand-off in between, so long-range connections fade.

Figure 1 · Interactive

The relay race and the round table

Reading it: the top lane is a recurrent network. Its memory (the glowing pill) visits one word per step, so the last word waits for all the others. The bottom lane is the Transformer: every word is processed in the same step, and the last word reaches the first in one hop. Drag the sentence length and watch the gap grow: the recurrent network's steps grow with the sentence, while the Transformer's stay at one per layer.

Others had tried to escape the relay. Convolutional models read words through small overlapping windows, which run in parallel, but two distant words only meet after many stacked layers. The Transformer's bet was more radical: drop the step-by-step memory entirely, and let every word connect to every other word directly.

“This inherently sequential nature precludes parallelization within training examples.”Vaswani et al. (2017), §1

Chapter 2 Paper §3.2 ↗

Every word looks at every word

You walk into a library and type a search: your query. Every book has a catalogue card describing it, its key, and the book itself on the shelf, its value. The librarian compares your search with every card and hands you not one book but a blend: mostly the best match, a little of the next best.

In a Transformer, every word plays all three parts at once. Each word searches, is catalogued and gets borrowed. This is self-attention, and it's the only place in the whole model where words exchange information.

“An attention function can be described as mapping a query and a set of key-value pairs to an output.”Vaswani et al. (2017), §3.2

Chapter 3 Paper §3.2.1 ↗

One attention step, by hand

Let's do the whole calculation for the word it, with numbers small enough to check in your head. Each word gets three short lists of numbers (vectors): a query, a key and a value. Real models learn these, and use 64 numbers each. We'll use four, and two for the values so we can draw them.

The formula, decoded

In words: score every query against every key, shrink the scores by the square root of the key length, turn each row into shares, and use the shares to blend the values.

The same numbers in Python
import math
q = [1, 1, 1, 1]                                   # the query for "it"
keys = {"animal": [1, 1, 1, 1], "tired": [1, 1, 0, 0], "street": [1, 0, 0, 0]}
values = {"animal": [1, 0], "tired": [0, 1], "street": [1, 1]}

scores = {w: sum(a * b for a, b in zip(q, k)) / math.sqrt(len(q)) for w, k in keys.items()}
total = sum(math.exp(s) for s in scores.values())
shares = {w: math.exp(s) / total for w, s in scores.items()}
# {'animal': 0.63, 'tired': 0.23, 'street': 0.14}

out = [sum(shares[w] * values[w][i] for w in values) for i in range(2)]
# [0.77, 0.37]

Footnote: why divide by √dk?

Roll one die and the result swings between 1 and 6. Add up a hundred dice and the total swings by dozens. A dot product is a sum of dk small products, so the longer the vectors, the wilder the scores. Wild scores make softmax hand nearly everything to one word, and then it stops responding to small changes. The gradient that training depends on vanishes.

Figure 3 · Interactive · computed live

Softmax with and without the scaling

Reading it: each panel shows one query's shares over eight random keys, with every number drawn from a standard normal distribution. On the left the raw dot products go straight into softmax; on the right they're divided by √dk first. Push the key length up: the left panel collapses onto a single key, while the right keeps a spread of shares that training can still adjust. The paper used dk = 64.
“We suspect that for large values of dk, the dot products grow large in magnitude, pushing the softmax function into regions where it has extremely small gradients.”Vaswani et al. (2017), §3.2.1

Chapter 4 Paper §3.2.2 ↗

Many heads, many questions

One attention can only ask one kind of question at a time. So give the same sentence to eight readers, each with a different coloured highlighter. One marks who did what, another which pronoun points where, another the word just before. Afterwards, staple their notes together. That's multi-head attention: eight small attentions running side by side, each free to learn its own question.

Figure 4 · Gallery

Six heads, six habits

Rows and columns run through the same eight words: The animal didn't cross because it was tired. Hover a cell to read it.

Reading it: each grid is one head. Rows are the word doing the looking; columns are the word looked at; brighter means more attention. Hover a cell to read it. The patterns are drawn by hand, but they're the kinds of habits researchers have found in trained models: heads that track the previous word, heads that link a pronoun to its noun, and heads that park their attention on the first word when there's nothing useful to look at (see attention sinks).

Splitting 512 into 8 × 64

The paper's word vectors hold 512 numbers. Rather than one attention over all 512, it runs 8 heads over 64 numbers each, glues the eight results back together, and mixes them with one more matrix. The total work is about the same as a single full-width head.

6464646464646464

In words: for each head, project the queries, keys and values down to 64 numbers with that head's own matrices and run attention; then place the eight outputs side by side and mix them with one final matrix WO.

Counting weights: each of WQ, WK, WV is 8 heads × (512 × 64), and WO is 512 × 512, so one attention layer holds 4 × 512 × 512 = 1,048,576 projection weights.

The paper's Table 3 tested the trade-off. With the total width fixed, a single head scored 0.9 BLEU worse than eight, and 32 narrow heads also did worse: more heads help, but only up to a point. Chapter 10 charts it.

“Multi-head attention allows the model to jointly attend to information from different representation subspaces at different positions.”Vaswani et al. (2017), §3.2.2

Chapter 5 Paper §3.2.3 ↗

No peeking

The translation model has two halves. The encoder reads the whole English sentence and writes rich notes about every word. The decoder writes the German sentence one word at a time. The same look-around-and-blend tool is used in three places, differing only in who asks and who answers:

1 · Encoder self-attention

The reader rereads their own notes. Every English word sees every other English word.

2 · Masked decoder self-attention

The writer rereads their draft, but never peeks at words not yet written.

3 · Encoder–decoder attention

The writer consults the reader's notes: German queries, English keys and values.

During training the model sees the whole correct German sentence at once, which is what makes training parallel. To stop the decoder cheating by reading the word it's meant to predict, the future is blanked out with a causal mask.

Figure 5 · Interactive

The causal mask

Hover or tap a cell.

Reading it: each row is one German word looking back at the sentence. Lit cells are allowed; dark ones are set to minus infinity before softmax, and e−∞ = 0, so they get exactly zero share. The allowed cells form a triangle: the first word sees only itself, the last sees everything before it. Chat models such as GPT and Claude are built from this one kind of attention alone: they're decoder-only.

Chapter 6 Paper §3.5 ↗

Where am I?

Attention has a blind spot: it has no sense of order. “Dog bites man” and “man bites dog” contain the same words, so without help they'd get the same result. The fix is to stamp each word with its position before anything else happens.

The paper's stamp is a wall of clocks whose hands turn at different speeds: a hand that spins every few words, one that takes dozens, one that takes thousands. Read all the hands together and you know exactly where you are. Better still, moving three words along always turns each hand by the same amount, which makes “three words later” easy for the model to recognise.

Figure 6 · Interactive · computed live

The wall of clocks

dimension 0 (fast) →dimension 127 (slow)

Reading it: each dial is one pair of numbers in the 512-number stamp, drawn as a hand at angle position ÷ 100002i/512. Drag the position: the left dials spin, the right ones barely move. The heatmap below shows the same stamps for positions 0 to 99 (rows) across the first 128 numbers (columns), with the current position outlined. Read any row across and you get a barcode unique to that position.

In words: for position pos, each pair of numbers gets the sine and cosine of the position divided by a number that grows from 1 to 10,000 across the pairs. Early pairs spin fast; late pairs spin slowly. The stamp is added to the word's embedding.

The paper also tried learned position vectors and got nearly identical results (Table 3, row E: 25.7 BLEU against 25.8). Most modern models use RoPE instead, which rotates queries and keys by an angle set by position: the same clock-hand idea, applied inside attention.

Chapter 7 Paper §3 ↗

The whole machine

Now assemble the parts. Words become vectors, positions are stamped on, and then the same layer repeats six times: attention, where words talk to each other, then a small feed-forward network, where each word thinks privately about what it heard. Around every step runs a shortcut, the residual connection: each layer writes corrections in the margin rather than rewriting the page. Then layer normalization tidies the numbers so nothing grows out of control.

Figure 7 · The paper's Figure 1, redrawn · Interactive

Follow a sentence through the Transformer

Scroll the diagram sideways to see the decoder →

Start here

Press Follow the data

Or pick any block. We'll translate “The cat sat.” into German with the paper's base model: vectors of 512 numbers, 8 heads, 6 layers.

Reading it: data flows from the bottom up. The encoder (left) runs its layer six times, shown by the dashed frame and the ×6. Its final notes travel along the long wire into the decoder's middle attention block. The decoder (right) climbs the same way over the German written so far, and at the top, Linear and Softmax turn the final vector into a probability for every word in the vocabulary. The arrows that jump around a block and rejoin at Add & Norm are the residual connections. Redrawn from Vaswani et al. (2017), Figure 1.

The two pieces that aren't attention

Feed-forward: widen each word's 512 numbers to 2,048, zero out the negatives (ReLU), and narrow back to 512. The same recipe runs on every word separately.

Add & norm: add the step's proposed change to its input, then rescale to zero mean and unit spread. With x = (2, 0) and a proposed change of (1, 1): (3, 1), which normalizes to (1, −1).

Per layer, the feed-forward network holds 2 × 512 × 2,048 + 2,048 + 512 = 2,099,712 parameters, about twice the attention projections. Across a Transformer, feed-forward layers hold roughly two thirds of the weights. The paper also shares one embedding matrix between the input, the output and the final linear layer, multiplying the embeddings by √512 ≈ 22.6 on the way in.

Chapter 8 Paper §4 ↗

Why it was faster

How many phone calls does it take for news to travel between two people? Along a chain of friends, as many calls as there are people between them. On a group call, one. The paper compares layer types on exactly this, plus how much work each does.

Table 1, reproduced with attribution (Vaswani et al., 2017). n is the sentence length, d the vector width, k the convolution window.
LayerWork per layerSteps in sequenceHops between two words
Self-attentionO(n²·d)O(1)O(1)
RecurrentO(n·d²)O(n)O(n)
ConvolutionalO(k·n·d²)O(1)O(logk n)

Figure 8 · Interactive · computed live

When is n² cheaper than d²?

Reading it: the bars compare the work in one layer, on a log scale, with d = 512 as in the paper. While sentences are shorter than the vector width, self-attention does less work, and in 2017 most sentences ran a few dozen tokens. Past n = 512 the balance flips. Today's contexts run to hundreds of thousands of tokens, which is why so much engineering since (FlashAttention, sliding windows, the KV cache) goes into taming the n² term.

Chapter 9 Paper §5 ↗

Training it

4.5MEnglish–German sentence pairs, split into about 37,000 sub-word tokens with byte-pair encoding
8NVIDIA P100 GPUs in one machine
12 hto train the base model: 100,000 steps at 0.4 seconds each
3.5 dto train the big model: 300,000 steps

A new driver pulls out of the driveway slowly, speeds up on the open road, then eases off near the destination. The learning rate follows the same shape. It starts tiny because an untrained model's gradients point in wild directions, climbs during a warmup, then decays so the model can settle into fine detail. The optimizer is Adam.

Figure 9 · Interactive · computed live

The warmup schedule

Hover the curve to read it.

Reading it: the curve rises in a straight line to a peak exactly at the end of warmup, then falls away as one over the square root of the step. The paper used 4,000 warmup steps, which peaks at about 7.0 × 10−4. A longer warmup gives a lower, later peak. After the peak, every setting follows the same decay.

In words: for the first warmup steps, grow the learning rate in a straight line; after that, shrink it as one over the square root of the step number; scale everything by one over the square root of the model width.

Two more tricks fight overfitting. Dropout (P = 0.1) zeroes a random tenth of each sub-layer's output during training. Label smoothing (ε = 0.1) marks the right answer as 90% right rather than 100%. The paper notes that label smoothing makes perplexity worse, because the model learns to be less sure, but improves BLEU: a training metric and a quality metric disagreeing.

Chapter 10 Paper §6 ↗

Better, and cheaper

BLEU scores a translation by how many of its word sequences also appear in a professional translation, on a 0 to 100 scale. A gain of one or two points counts as a clear step forward. The Transformer didn't just win; it won while costing less to train.

Figure 10 · Table 2, charted

Quality against training cost, English→German

Reading it: each dot is one model. Up is better translation (BLEU on newstest2014); left is cheaper training (floating-point operations, log scale). The base Transformer sits above every earlier model while costing about a third of ConvS2S's compute. The big Transformer spends as much as Google's GNMT system and beats it by nearly 4 points. Selected rows of Table 2, reproduced with attribution (Vaswani et al., 2017).

Figure 11 · Table 3 (A), charted

How many heads?

Reading it: BLEU on the development set (newstest2013) as the number of heads changes, with the total width held at 512 so more heads means narrower heads. One head is clearly worst; 8 and 16 tie at the top; 32 narrow heads fall back. Selected rows of Table 3, reproduced with attribution (Vaswani et al., 2017).

Chapter 11 Paper §7 ↗

What happened next

The paper closes by planning to take attention to images, audio and video, to study attention restricted to local windows for long inputs, and to make generation less sequential. All three happened. The core (attention and feed-forward layers wrapped in residual connections) is unchanged to this day. Nearly every detail around it has been tuned.

What changed since 2017

In the paperCommon todayWhy
Encoder + decoderDecoder-only for chat; encoder-only for embeddingsOne stack of masked attention generates text; a two-way stack makes the best embeddings
Norm after each stepNorm before each step, often RMSNormMore stable training at depth
Sine and cosine positionsRoPERelative positions, and easier to stretch to longer contexts
ReLU feed-forwardGELU or SwiGLU, sometimes mixture of expertsBetter quality per parameter
8 heads, each with its own keys and valuesGrouped-query attentionA smaller KV cache, so more users per GPU
Plain attentionFlashAttention kernelsThe same math with far less memory traffic
6 layers, 65M to 213M parametersDozens of layers, billions of parametersScaling laws: bigger models on more data keep improving
“In this work, we presented the Transformer, the first sequence transduction model based entirely on attention.”Vaswani et al. (2017), §7

Test yourself

Five questions

Each one checks an idea, not a fact you could look up. Pick an answer to see why it's right or wrong.

Build it yourself

Every piece, in plain Python.

The AI Primer builds each idea on this page from scratch with NumPy alone, with tests that prove it works. Its lessons are being illustrated here, starting with attention.

About this page. This article explains the paper; it doesn't reproduce it. It quotes a sentence at most per chapter, clearly marked, and every figure is redrawn from scratch. The tables are reproduced in part with attribution under the notice printed on the paper: “Provided proper attribution is provided, Google hereby grants permission to reproduce the tables and figures in this paper solely for use in journalistic or scholarly works.” Numbers marked illustrative were chosen to teach; everything else comes from the paper or is computed live from its formulas. Read the original alongside: every chapter links to its section.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł. and Polosukhin, I. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems 30. arXiv:1706.03762.