rumblr Work in progressWIP

● The AI Primer · Lesson 2 · Part 1: how the model works inside

Neural networks

neurons, activations and backprop by hand

This lesson covers Neurons, activations, the forward pass, backprop by hand

Free lesson 30 min9 figures and diagrams7 interactive
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. A neuron is a weighted sum plus bias through a nonlinearity; a layer is a matrix multiply; a network is stacked layers.
  2. Without nonlinearity, depth is pointless: stacked linear layers equal one.
  3. Training loop: forward, loss, backprop (the chain rule gives a gradient per weight), optimizer step. Repeat.
  4. Backprop reuses forward activations, so training needs far more memory than inference.

Level 2

How it works, from scratch

A neuron is six lines of arithmetic, a layer is a matrix multiply, and training is a loop of four boxes. This level builds each one with numbers small enough to check by hand, then verifies the hand-written gradients against a numerical check.

Chapter 1

The idea: learning is adjusting knobs to be less wrong

Press Step three times before reading: those are the three rows of the lesson's table.

Picture learning to throw darts blindfolded, with a friend calling out "a bit high, a bit left". You throw, hear how far off you were, and adjust your arm a little. After enough throws your arm "knows" where the bullseye is. A neural network learns the same way. Its "arm" is millions of numbers called weights, and the friend is a formula (the loss) that says how far off each guess was and which way to adjust.

Two words carry this whole lesson. The derivative (or slope) of the loss with respect to a weight answers "if I nudge this weight up a tiny bit, does the loss go up or down, and how fast?" The gradient is simply the list of those slopes, one per weight. Stepping against the gradient is walking downhill. (New to the notation? primer.notation builds every symbol used here from zero.)

Worked example with a single knob: a model with one weight w predicts y = w·x and should learn y = 3x from the example x = 1, y = 3. The loss is (w·1 − 3)². Its slope with respect to w is 2(w − 3). Start at w = 0 and step against the slope with step size 0.25:

step w slope 2(w − 3) new w = w − 0.25·slope
0 0 −6 1.5
1 1.5 −3 2.25
2 2.25 −1.5 2.625

Each step halves the distance to 3. That's all training is.

Figure 1 · Diagram

Reading it: follow the arrows clockwise. A batch of examples goes in; the forward pass turns them into predictions; the loss scores how wrong those predictions were as a single number; backpropagation works out, for every weight, which direction would have made that number smaller; the optimizer nudges each weight a small step that way. Then the next batch arrives. The single-knob table above is one trip around this loop, three times.

The update rule, for every weight at once, with learning rate η:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
one weight (a knob the model can turn) 0 at the start
"replace the old value with" (an update, not an equation)
"eta", the learning rate: how big a step to take 0.25
the loss: one number saying how wrong the model is
the slope of the loss with respect to (the "partial derivative": how the loss changes when only moves)

In words: "the new weight is the old weight minus the learning rate times the slope of the loss at the old weight."

With the numbers: , the first row of the table.

Level 3: in Python
w, eta, target = 0.0, 0.25, 3.0
for step in range(3):
    # ∂L/∂w for the loss (w - 3)²
    slope = 2 * (w - target)
    # w ← w - η ∂L/∂w
    w = w - eta * slope
    print(w)  # → 1.5 2.25 2.625

learn_one_weight runs the table above; train runs the loop for a real network.

Why it matters this loop is identical for a spam filter and for a frontier language model. Only the size of the network, the data, and the loss change. When a training run misbehaves, the fault is in one of these four boxes.

Chapter 2

A neuron: a weighted vote followed by a decision

Think of a judge at a cooking contest. They score taste, presentation and originality, but care about taste most, so they weight it more. They add a baseline ("everyone starts at half a point"), then apply a rule to turn the total into a verdict ("negative totals count as zero"). That is a neuron: weights are how much each input matters, the bias is the baseline, and the activation function is the rule.

Worked example: inputs (2, 3), weights (0.5, 1.0), bias 0.5. Weighted sum: 2 × 0.5 + 3 × 1.0 + 0.5 = 4.5. The ReLU rule keeps positive values, so the output is 4.5.

Figure 2 · Diagram

Reading it: each input travels along an edge that multiplies it by that edge's weight; the middle node adds the products and the bias; the activation reshapes the sum into the output. A layer is this picture repeated side by side, every input wired to every neuron, which is exactly a matrix multiply. A network is layers stacked.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the -th input ,
the weight on the -th input ,
a counter over the inputs 1, 2
"add up the following for every "
the bias (the baseline added to every sum) 0.5
the dot product: multiply matching entries of two lists, then add
"phi", the activation function ReLU
the neuron's output 4.5
a batch of inputs, one example per row shape (batch, inputs)
a layer's weights, one column per neuron shape (inputs, neurons)
matrix multiply: every row of dotted with every column of shape (batch, neurons)
every neuron's output for every example shape (batch, neurons)

In words: "multiply each input by its weight, add them up, add the bias, and pass the total through the activation function."

With the numbers: $y = \text{ReLU}(0.5 \cdot 2 + 1.0 \cdot 3 + 0.5) = \text{ReLU}(4.5) = 4.5$.

Level 3: in Python
x = [2, 3]
w = [0.5, 1.0]
b = 0.5
# ReLU: keep positives, zero the rest
def phi(z): return max(0.0, z)
# w · x = Σ_i w_i x_i
sum(w_i * x_i for w_i, x_i in zip(w, x))  # → 4.0
# y = φ(w · x + b)
phi(sum(w_i * x_i for w_i, x_i in zip(w, x)) + b)  # → 4.5

Why it matters every dense layer in every model is this, including the feed-forward half of each transformer block, where most of a language model's parameters live. MLP.forward is two of these layers in six lines.

Chapter 3

Backpropagation: passing the blame backwards

Take the step yourself first; the lesson's table below is the same step, written out.

Picture an assembly line that produced a faulty product. The inspector at the end measures how bad it is and passes the complaint back. Each station works out how much of the fault was its own doing, based on what it received and what it did to it, and passes the rest further back. Every station ends up knowing exactly how to adjust its own machine. Backpropagation is that blame-passing, done with derivatives.

Worked example, continuing the neuron: the target was 5, so the loss is (5 − 4.5)² = 0.25.

link local derivative running blame
loss → y 2(y − 5) = −1 −1
y → z (ReLU, z > 0) 1 −1
z → w1 x1 = 2 −2
z → w2 x2 = 3 −3
z → b 1 −1

One step with learning rate 0.01 moves weights to (0.52, 1.03) and bias to 0.51. The new output is 4.64 and the loss drops from 0.25 to 0.1296. The weight whose input was bigger (3) received the bigger share of blame and moved more.

Figure 3 · Diagram

Reading it: this is the neuron read right to left. Start at the loss and multiply the local derivatives along each path; that product is the chain rule. All paths share the first two links (−1 × 1), then split. Each weight's gradient is the blame arriving at the sum times the input that weight multiplied. Negative gradient means "increase me to reduce the loss".
Level 3: the formula and its symbols

This is the chain rule: when a change passes through several steps, the overall rate of change is the product of each step's rate of change. If turning a knob moves a gear 3× as fast, and that gear moves a needle 2× as fast, the knob moves the needle 6× as fast.

Symbols

Symbol Meaning here In the example
the loss 0.25
the target (the right answer) 5
the neuron's output 4.5
the weighted sum before the activation 4.5
how the loss changes as the output changes
how the output changes as the sum changes: the activation's slope ReLU slope at 4.5 = 1
how the sum changes as weight changes: just its input
the activation's derivative (slope) 1 for ReLU when

In words: "the blame on a weight is how much the loss cares about the output, times how much the output cares about the sum, times how much the sum cares about that weight."

With the numbers: for : ; for : .

In Python:

x, t, z = [2, 3], 5, 4.5
# ReLU(4.5)
y = max(0.0, z)
# ∂L/∂y
dL_dy = 2 * (y - t)
# φ'(z): ReLU's slope
dy_dz = 1 if z > 0 else 0
# × ∂z/∂w_i = x_i, for w_1 and w_2
[dL_dy * dy_dz * x_i for x_i in x]  # → [-2.0, -3.0]
# one step, learning rate 0.01
w = [0.5 - 0.01 * -2.0, 1.0 - 0.01 * -3.0]
b = 0.5 - 0.01 * -1.0
y_new = w[0] * x[0] + w[1] * x[1] + b
# the new output and the smaller loss
round(y_new, 2), round((t - y_new) ** 2, 4)  # → (4.64, 0.1296)

one_neuron_worked_example computes every row of the table.

Why it matters PyTorch's loss.backward() does exactly this, for billions of weights, by recording the forward computation and replaying it backwards. Knowing it is just the chain rule is what lets you reason about vanishing gradients, memory use, and why some architectures train and others don't.

Chapter 4

Why the nonlinearity is essential

Stack three sheets of tinted glass and you get one darker sheet of glass: no arrangement of flat panes can bend light around a corner. Linear layers are flat panes. However many you stack, the result is one linear layer, which can only separate classes with a straight line. The activation function is the bend.

Worked example in one dimension: a layer that multiplies by 3, followed by a layer that multiplies by 2, is the same as one layer that multiplies by 6. With a ReLU between them, inputs −1 and 1 give 0 and 6. No single multiplication does that.

Figure 4 · Diagram

Reading it: in the top row the multiplications are chained with nothing in between, and associativity lets you multiply the weights together first: three layers equal one. In the bottom row a nonlinearity φ sits between the multiplies, the product can no longer be pre-computed, and the network can bend its decision boundary.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the inputs or
two layers' weights 3 and 2
the two layers multiplied into one 6
an activation between the layers ReLU
any single weight you might try instead none works
"is not equal to"

In words: "two linear layers in a row are the same as one linear layer, but put an activation between them and no single layer can copy them."

With the numbers: without ReLU, and , exactly . With ReLU, and . A single weight would need and at once: impossible.

Level 3: in Python
W_1, W_2 = 3, 2
# ReLU
def phi(z): return max(0, z)
# (X W_1) W_2 ...
[(x * W_1) * W_2 for x in [-1, 1]]  # → [-6, 6]
# ... equals X (W_1 W_2): one layer of 6
[x * (W_1 * W_2) for x in [-1, 1]]  # → [-6, 6]
# φ(X W_1) W_2: no single W' gives 0 and 6
[phi(x * W_1) * W_2 for x in [-1, 1]]  # → [0, 6]

Figure 5 · Chart

−1.5 −1.0 −0.5 0.0 0.5 1.0 1.5 2.0 2.5 x1 −1.0 −0.5 0.0 0.5 1.0 1.5 x2 Linear (logistic regression): 88% class 0 class 1 −1.5 −1.0 −0.5 0.0 0.5 1.0 1.5 2.0 2.5 x1 MLP, 16 tanh hidden units: 99.75%

Two-moons decision boundaries: logistic regression's straight line misclassifies the moon tips (88%), while one hidden tanh layer bends around the gap (99.75%)

Reading it: both panels show the same two interleaving half-moons, coloured by class; the shaded background is what each model predicts at every point. On the left, a single linear layer can only split the plane with a straight line, so the tips of both moons land on the wrong side (about 88% accuracy). On the right, one hidden tanh layer bends the boundary to follow the gap between the moons (99.75%, one point of 400 wrong). Same data, same optimizer: the only difference is a nonlinearity between two layers.

In code: make_moons builds the two interleaving half-moons (and make_xor the four-point XOR puzzle); train_logistic_regression fits the straight-line baseline in the left panel.

Why it matters "depth" is only worth anything because of the activations between layers. linear_stack_collapses shows five stacked linear layers matching their single-matrix product to 1e-16.

The two-moons figure above, trained in front of you, with a switch for the bend.

Chapter 5

Activation functions: the rule after the sum

Think of different kinds of switch. ReLU is a one-way valve: water flows freely forward, not at all backward. Sigmoid is a dimmer that squeezes any input into 0–1 but barely moves at the extremes. Tanh is the same dimmer centred on zero. GELU is a valve that leaks a little when nearly closed.

Worked example, values you can check on a calculator (the ′ mark, as in sigmoid′, means "the slope of"):

z ReLU sigmoid sigmoid' tanh' GELU
−1 0 0.269 0.197 0.420 −0.159
0 0 0.5 0.25 (its maximum) 1 (its maximum) 0
1 1 0.731 0.197 0.420 0.841

Figure 6 · Chart

−4 −2 0 2 4 input z −1.5 −1.0 −0.5 0.0 0.5 1.0 1.5 2.0 2.5 3.0 output φ(z) Activation functions relu gelu sigmoid tanh −4 −2 0 2 4 input z 0.0 0.2 0.4 0.6 0.8 1.0 derivative φ'(z) Their derivatives (gradient that flows back) relu gelu sigmoid tanh

Activations and their slopes: sigmoid's slope peaks at 0.25 and sigmoid and tanh slopes vanish past |z| = 3, while ReLU's slope is a clean step from 0 to 1

Reading it: the left panel shows each activation's output, the right its derivative, which is how much gradient passes back through it. Look at the tails of the right panel first: sigmoid and tanh derivatives fall to almost zero once |z| passes 3, so a saturated neuron barely learns. Sigmoid's derivative never exceeds 0.25 even at its peak; stack ten such layers and the gradient can shrink by 0.25¹⁰ ≈ 1e-6. ReLU's derivative is a step: exactly 1 for positive inputs (no shrinking) and exactly 0 for negatives. GELU's is a smooth version of that step.
Function Formula Range Where it's used
ReLU max(0, z) [0, ∞) Hidden layers of CNNs and MLPs. Fast. A neuron stuck negative gets zero gradient forever ("dying ReLU").
GELU z·Φ(z), Φ = normal CDF ≈[−0.17, ∞) Transformers (BERT, GPT-2 and most since).
Sigmoid 1 / (1 + e^−z) (0, 1) Binary outputs, and the gates inside LSTMs.
Tanh (e^z − e^−z)/(e^z + e^−z) (−1, 1) RNN hidden states; zero-centred.
Softmax e^{z_i} / Σ e^{z_j} probabilities The output layer over classes or tokens, and inside attention.

Reading the formulas: e is Euler's number, about 2.718, and e^z means "2.718 raised to the power z", which is always positive and grows fast. Φ(z) is the fraction of a bell curve (the standard normal distribution) that lies below z: Φ(0) = 0.5, Φ(1) = 0.841. Softmax turns a list of scores into shares that are all positive and add up to 1: raise e to each score, then divide each by the total (Σ means "add them all up"). See primer.notation for each symbol, and primer.ml.attention for softmax worked through by hand.

In code: relu, sigmoid, tanh and gelu compute each rule, and relu_grad, sigmoid_grad, tanh_grad and gelu_grad compute its slope; gelu_tanh is the cheaper approximation GPT-2 uses, and softmax subtracts the largest score first so nothing overflows.

Why it matters activation choice decides whether gradients survive a deep stack. Sigmoid everywhere is why deep networks were hard to train before 2010; ReLU and residual connections are why they aren't now (see primer.ml.deep_nets).

Chapter 6

Backprop through two layers

Run the nine lines one at a time; the lesson's table below is the same nine.

Now the assembly line has two stations. The blame from the end first reaches the output weights, then travels back through the hidden layer to the first weights. On the way it passes through the hidden activation, which can shrink it, just as a station that barely changed the product can take only a little of the blame.

Worked example, one input, one hidden unit, one output: x = 1, w1 = 0.5, tanh, w2 = 2, sigmoid, target 1.

step value
z1 = x·w1 0.5
h = tanh(0.5) 0.4621
z2 = h·w2 0.9242
ŷ = σ(0.9242) 0.7159
dL/dz2 = ŷ − y −0.2841
dL/dw2 = h · dL/dz2 −0.1313
dL/dh = dL/dz2 · w2 −0.5682
dL/dz1 = dL/dh · (1 − h²) −0.5682 × 0.7864 = −0.4469
dL/dw1 = x · dL/dz1 −0.4469

Figure 7 · Diagram

Reading it: the top row is the forward pass, left to right; the dotted arrows are the backward pass, starting at the loss and walking back one box at a time. Each label is one line of MLP.backward, and the worked table above is that same path with numbers. Notice the three "cached" arrows: the backward pass reuses h and X from the forward pass. That dependency is why training must keep every layer's activations in memory, and inference does not.

For a batch, with binary cross-entropy (the loss for yes/no predictions, ; see primer.ml.losses), which fuses with sigmoid into the simple error ŷ − y:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example (one example, one unit)
inputs, one row per example
first and second layer weights 0.5, 2
each layer's weighted sum, before its activation 0.5, 0.9242
hidden activations, 0.4621
"y-hat", the prediction 0.7159
the target 1
sigmoid, squashes any number into 0 to 1
hyperbolic tangent, squashes into −1 to 1
"transpose": flip a matrix so rows become columns, so the shapes line up for the multiply a scalar is its own transpose
multiply element by element (not a matrix multiply)
the slope of tanh at 0.7864

In words: "the error at the output is prediction minus target; each layer's weight gradient is that layer's input times the error arriving at it; to send the error one layer further back, multiply by the weights it came through and by the activation's slope."

With the numbers: ; ; ; ; .

Level 3: in Python
import math
X, W_1, W_2, y = 1, 0.5, 2, 1
z_1 = X * W_1
h = math.tanh(z_1)
z_2 = h * W_2
# σ(z_2)
y_hat = 1 / (1 + math.exp(-z_2))
# ŷ - y
dL_dz2 = y_hat - y
# hᵀ ∂L/∂z_2
dL_dW2 = h * dL_dz2
# ∂L/∂z_2 W_2ᵀ
dL_dh = dL_dz2 * W_2
# ⊙ (1 - h²), tanh's slope
dL_dz1 = dL_dh * (1 - h ** 2)
# Xᵀ ∂L/∂z_1
dL_dW1 = X * dL_dz1
[round(v, 4) for v in (dL_dz2, dL_dW2, dL_dh, dL_dz1, dL_dW1)]  # → [-0.2841, -0.1313, -0.5682, -0.4469, -0.4469]

The hand-written gradients are checked two ways: a numerical gradient check (nudge each weight by ±ε and measure the loss change; see gradient_check) and, in the tests, PyTorch autograd.

In code: tiny_two_layer_example computes every row of the worked table; MLP holds both layers' weights and biases, MLP.loss scores a batch with binary_cross_entropy, and numerical_gradient is the slow nudge-every-weight answer that gradient_check compares against.

Why it matters the factor (1 − h²) is where gradients shrink. Chain fifty of them and the first layers hear almost nothing: the vanishing gradient problem. The cached activations are why training a model needs several times the memory of serving it.

Chapter 7

Batch, step, epoch

Imagine studying a deck of 400 flashcards. You work through a pile of 32, then pause to update your notes; that pause is a step. The pile is a batch. Going through the whole deck once is an epoch; then you shuffle and start again.

Worked example: 400 examples in batches of 32 is 12 full batches plus one of 16, so ⌈400/32⌉ = 13 steps per epoch. Two epochs: 26 steps.

Figure 8 · Diagram

Reading it: an epoch reshuffles the data and slices it into batches; every batch produces exactly one weight update. Shuffling each epoch means batches differ every time, so the gradient noise doesn't repeat.

Figure 9 · Chart

1 0 0 1 0 1 1 0 2 epoch (log scale) 0.00 0.05 0.10 0.15 0.20 0.25 0.30 training loss (BCE) Training curve on two moons

Training loss per epoch on a log axis: slow at first, a steep fall once hidden units find features, then a flat tail near zero

Reading it: the horizontal axis counts epochs on a log scale; the vertical axis is the loss on the whole training set after each one. Loss drops slowly at first while the weights are near their random start, falls steeply once the hidden units find useful features, then flattens as the remaining errors get harder. Plotting this curve is the first thing to do when training anything.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
number of training examples 400
batch size 32
"ceiling": round up to the next whole number (the last, smaller batch still counts as a step)

In words: "an epoch takes as many steps as it takes batches of size B to cover N examples, rounding up."

With the numbers: steps per epoch; 2 epochs = 26 steps.

Level 3: in Python
import math
N, B, epochs = 400, 32, 2
N / B  # → 12.5
# ⌈N / B⌉: the last, smaller batch still counts
math.ceil(N / B)  # → 13
epochs * math.ceil(N / B)  # → 26

In code: train reshuffles the data every epoch, takes one plain gradient step per batch and records the loss after each epoch; MLP.accuracy reports the fraction of examples classified correctly.

Why it matters larger batches give smoother gradient estimates and keep GPUs busy, but need more memory. Pretraining a language model runs roughly one epoch over trillions of tokens (it rarely sees the same text twice); fine-tuning runs several epochs over a small dataset, which is why fine-tunes can overfit.

Test yourself

6 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Why can't you just stack linear layers?Think it through, then reveal

Their composition is a single linear map (W1·W2 is just another matrix), so extra layers add parameters but no expressive power.

Question 2What does backprop actually compute?Think it through, then reveal

The gradient: for every weight, how much a tiny change would change the loss. It applies the chain rule from the output backward, reusing values cached in the forward pass.

Question 3Why is ReLU preferred over sigmoid in hidden layers?Think it through, then reveal

Sigmoid's derivative is at most 0.25 and near 0 when saturated, so gradients shrink multiplicatively with depth. ReLU's derivative is exactly 1 for positive inputs. The cost: neurons whose input stays negative get zero gradient and "die".

Question 4Why does training need more memory than inference?Think it through, then reveal

The backward pass needs every layer's activations from the forward pass, so they must be kept until the gradients are computed. Inference can discard each activation as soon as the next layer has used it.

Question 5What's the gradient of sigmoid + binary cross-entropy with respect to the logit?Think it through, then reveal

ŷ − y: prediction minus target. Clean, bounded and cheap, which is why the two are always fused.

Question 6Batch vs. step vs. epoch?Think it through, then reveal

A batch is the examples per update, a step is one update, an epoch is one full pass over the data.

Primary sources

The papers behind this lesson

Rumelhart, Hinton & Williams, Learning representations by back-propagating errors (Nature, 1986): Showed that the chain rule, run backwards through a multi-layer network, trains hidden units to discover useful internal features.

The paper ↗

Hendrycks & Gimpel, Gaussian Error Linear Units (GELUs) (2016): Introduced GELU, the smooth ReLU that became the default activation in transformers.

The paper ↗

Researcher's shelf

Further reading

  • CS231n notes, Backpropagation, intuitions: https://cs231n.github.io/optimization-2/
  • CS231n notes, Neural Networks Part 1: https://cs231n.github.io/neural-networks-1/
  • Andrej Karpathy, micrograd (autograd in about 100 lines): https://github.com/karpathy/micrograd
  • Andrej Karpathy, The spelled-out intro to neural networks and backpropagation: https://www.youtube.com/watch?v=VMj-3S1tku0
  • 3Blue1Brown, Neural networks series: https://www.3blue1brown.com/topics/neural-networks
  • Michael Nielsen, Neural Networks and Deep Learning, ch. 2: http://neuralnetworksanddeeplearning.com/chap2.html
  • Hendrycks & Gimpel, GELU (2016): https://arxiv.org/abs/1606.08415
  • PyTorch autograd tutorial: https://pytorch.org/tutorials/beginner/blitz/autograd_tutorial.html

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.