At a glance
Key takeaways
- Overfitting: training error falls while validation error rises; the model memorised noise. Underfitting: both errors are high.
- Fixes for overfitting: more data, early stopping, dropout, weight decay (L2), L1 for sparsity, data augmentation.
- Bias is systematic error (too simple); variance is sensitivity to the training sample (too flexible).
- Train fits weights, validation tunes choices, test is touched once; cross-validation for small data.
- Leakage (duplicates across splits, features from the future) gives great offline scores that collapse in production.
Level 2
How it works, from scratch
Two students prepare for an exam. One learns the ideas. The other memorises last year's paper, answer by answer, and scores 100% on it in practice. On the real exam, with new questions, the first student does fine and the second falls apart. A model that memorises its training examples, noise and all, instead of learning the pattern behind them is overfitting. It looks brilliant on the data it has seen and fails on data it hasn't. Regularization is everything we do to push a model toward the first student.
Worked example: four training points (0, 0), (1, 1), (2, 0), (3, 1), a zig-zag that is really "about 0.5, plus noise".
| model | error on the 4 training points | prediction at x = 4 |
|---|---|---|
| flat line at 0.5 (the pattern) | 0.25 | 0.5 |
| cubic through all four points (memorised) | 0 | 8 |
The cubic is perfect on the training data and absurd one step beyond it. You can check the 8 by hand with a difference table: the values 0, 1, 0, 1 have differences 1, −1, 1, then −2, 2, then 4; a cubic keeps that last difference constant, so extending the table gives 6, then 7, then 8.
Figure 1 · Diagram
flowchart LR
D["Training data =<br/>pattern + noise"] --> M{Model capacity}
M -->|too little| U["Underfits:<br/>misses the pattern"]
M -->|about right| G["Generalizes:<br/>learns the pattern"]
M -->|too much, unchecked| O["Overfits:<br/>memorises the noise"]
G --> N[Good on new data]
U & O --> B[Bad on new data]
Figure 2 · Chart
Three polynomial fits to 12 noisy sine samples: the degree-1 line misses the curve, the cubic follows it, and the degree-11 polynomial hits every dot but swings wildly between them
The two numbers that tell these apart are errors on two different sets of data:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| how many training and validation examples | 4 training points | |
| "for every example in the training set" | ||
| the true value | 0, 1, 0, 1 | |
| the model's prediction | 0.5 everywhere for the flat line | |
mean squared error (see primer.ml.losses) |
0.25 for the flat line |
In words: "training error is the average squared miss on the data the model learned from; validation error is the same average on data it has never seen."
With the numbers: flat line: . Cubic: every miss is 0, so training error is 0; its validation error on any point beyond x = 3 is enormous.
Level 3: in Python
# the true values at x = 0, 1, 2, 3
y = [0, 1, 0, 1]
# the flat line predicts 0.5 everywhere
y_hat = [0.5, 0.5, 0.5, 0.5]
n = len(y)
# (1/n) Σ (y_i - ŷ_i)²
sum((y_i - y_hat_i) ** 2 for y_i, y_hat_i in zip(y, y_hat)) / n # → 0.25
In code: zigzag_fit_error and zigzag_predict fit a polynomial of any
degree to the four zig-zag points and report its training error and its
prediction; true_function and make_curve_data draw the noisy sine
samples in the figure.
Why it matters training error alone always rewards memorisation. The gap between training and validation error is the single most useful diagnostic in machine learning.
Chapter 1
Underfitting vs. overfitting: finding the middle
Goldilocks tries three bowls of porridge: too cold, too hot, just right. Model capacity works the same way, and you find "just right" by measuring, not guessing: sweep the capacity and watch both errors.
Worked example: 12 noisy points from a sine wave, 300 fresh points for validation, polynomials of increasing degree.
| degree | training error | validation error | diagnosis |
|---|---|---|---|
| 1 | 0.20 | 0.30 | underfit: both high |
| 3 | 0.021 | 0.065 | about right |
| 5 | 0.013 | 0.057 | best: lowest validation error |
| 9 | 0.004 | 0.71 | overfit: validation 12× worse than at degree 5 |
| 11 | 0.0000 | 7,631 | memorised |
Figure 3 · Chart
Error against polynomial degree on a log scale: training error only falls, while validation error is lowest at degree 5 and then climbs steeply
In code: polynomial_errors fits by least squares (choosing the
coefficients that minimise training MSE) and reports both errors.
Why it matters "both errors high" and "training low, validation high" need opposite fixes. Underfitting wants more capacity, more features or more training; overfitting wants more data, regularization or early stopping. Diagnosing which one you have is the point.
Chapter 2
Early stopping: take the cake out before it burns
A baker checks the cake with a toothpick every few minutes and takes it out when it comes out clean. Leave it longer and it burns, however good it was a minute ago. Training a flexible model is similar: validation error falls while the model learns the pattern, then rises once it starts memorising. Early stopping watches validation error during training, stops once it has failed to improve for a few checks (the patience), and keeps the weights from the best check.
Worked example: validation losses per epoch 1.0, 0.8, 0.6, 0.55, 0.58, 0.65, 0.7 with patience 2. Epoch 3 is the best. Epochs 4 and 5 are both worse, so training stops at epoch 5 and the weights from epoch 3 are kept.
Figure 4 · Diagram
flowchart LR
E[End of epoch] --> V[Measure validation loss]
V --> Q{Better than best?}
Q -->|yes| S[Save weights,<br/>reset counter]
Q -->|no| C[Counter + 1]
C --> P{Counter = patience?}
P -->|no| E2[Next epoch]
P -->|yes| R[Stop; restore<br/>best weights]
S --> E2
Figure 5 · Chart
Loss curves for an over-sized network: training loss keeps falling while validation loss bottoms out near epoch 300 and then rises
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the epoch number | 0 to 6 | |
| validation loss after epoch | 0.55 at | |
| "the at which the following is smallest" (the position, not the value) | 3 | |
| "t-star", the best epoch so far | 3 | |
| patience | how many non-improving epochs to tolerate | 2 |
In words: "remember the epoch with the lowest validation loss, and stop
once you've gone patience epochs past it without doing better."
With the numbers: ; at , , so stop.
Level 3: in Python
# validation loss after epoch t = 0, 1, 2, ...
L_val = [1.0, 0.8, 0.6, 0.55, 0.58, 0.65, 0.7]
patience = 2
t_star = 0
for t, loss in enumerate(L_val):
if loss < L_val[t_star]:
# a new best epoch: remember it
t_star = t
if t - t_star >= patience:
# patience used up: stop here
break
t_star, t # → (3, 5)
In code: early_stopping replays a run's validation losses and returns
the best epoch and the epoch where training stops; train_flexible_model
trains the over-sized network in the figure and records both losses.
Why it matters it's the cheapest regularizer there is: no change to the model, just a rule for when to stop. It's used almost everywhere a model is trained for multiple epochs, including fine-tuning language models.
Chapter 3
Bias and variance: two ways to miss the target
Picture two archers. The first groups every arrow tightly, but a hand's width to the left of the bullseye: consistently wrong. That's bias. The second's arrows are centred on the bullseye on average but scattered all over the target: inconsistent. That's variance. A model trained on a different sample of data is like another volley of arrows. Simple models behave like the first archer; very flexible ones like the second.
Worked example: fit polynomials to 300 different random training sets of 30 points each, and measure (averaged over the input range):
| degree | bias² (systematic error) | variance (swing between training sets) |
|---|---|---|
| 1 | 0.174 | 0.022 |
| 3 | 0.0035 | 0.015 |
| 9 | 0.0038 | 4.5 |
Figure 6 · Chart
Twenty fits from different training sets: degree-1 lines agree but all miss the sine (high bias), degree-9 curves average to the sine but scatter widely (high variance)
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the true pattern | ||
| "f-hat", the model's prediction, which depends on which training set it saw | one thin line | |
| expected value: the average over many random training sets (and noise) | average of the 300 fits | |
| a noisy observation, plus noise | ||
| the noise variance: error no model can remove |
In words: "a model's expected squared error on new data splits into how far its average prediction is from the truth (bias squared), plus how much its predictions scatter around their own average (variance), plus noise nobody can predict."
With the numbers: degree 1: ; degree 3: ; degree 9: . Degree 3 has the best total.
Level 3: in Python
sigma = 0.3
# σ²: the error no model can remove
noise = sigma ** 2
for degree, bias_sq, variance in [(1, 0.174, 0.022), (3, 0.0035, 0.015), (9, 0.0038, 4.5)]:
# bias² + variance + σ²
total = bias_sq + variance + noise
# two significant figures
print(degree, f"{total:.2g}") # → 1 0.29 3 0.11 9 4.6
In code: bias_variance fits one polynomial to each of many random
training sets and splits their error into bias², variance and noise.
Why it matters it names the two failure modes and explains why more data helps flexible models (averaging tames variance) but not rigid ones (more data doesn't fix bias). Very large neural networks complicate the picture ("double descent": past a point, even bigger models generalize better again), but the vocabulary is universal.
Chapter 4
Dropout: nobody gets to be indispensable
A coach who randomly sends half the team home before each practice forces every player to learn every position; no single star can carry the team. Dropout does this to neurons: during training, each activation is set to zero with probability p on every step, so the network can't rely on any single neuron or fragile combination of neurons.
Worked example: four activations (1, 1, 1, 1), drop rate p = 0.5. Suppose the random mask keeps the last two: (0, 0, 1, 1). The survivors are scaled by 1 / (1 − 0.5) = 2, giving (0, 0, 2, 2). The average is still 1, so the next layer sees the same size signal on average. At evaluation time, nothing is dropped and nothing is scaled.
Figure 7 · Diagram
flowchart LR
subgraph Train["Training step"]
a1((1)) --> k1((0))
a2((1)) --> k2((0))
a3((1)) --> k3((2))
a4((1)) --> k4((2))
end
subgraph Eval["Evaluation"]
b1((1)) --> e1((1))
b2((1)) --> e2((1))
b3((1)) --> e3((1))
b4((1)) --> e4((1))
end
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| a layer's activations | (1, 1, 1, 1) | |
| "h-tilde", the activations after dropout | (0, 0, 2, 2) | |
| the random mask of 0s and 1s | (0, 0, 1, 1) | |
| each mask entry is independently 1 with probability , else 0 (a coin flip) | 50/50 | |
| the drop rate | 0.5 | |
| multiply entry by entry |
In words: "flip a biased coin for each activation, zero the ones that lose, and divide the survivors by the keep probability."
With the numbers: .
Level 3: in Python
import random
random.seed(0)
h = [1, 1, 1, 1]
p = 0.5
# m_i ~ Bernoulli(1 - p): a coin flip each
m = [1 if random.random() < 1 - p else 0 for _ in h]
m # → [0, 0, 1, 1]
# (m ⊙ h) / (1 - p)
[m_i * h_i / (1 - p) for m_i, h_i in zip(m, h)] # → [0.0, 0.0, 2.0, 2.0]
In code: dropout draws the mask and scales the survivors during
training, and passes activations through unchanged at evaluation time.
Why it matters dropout was a key ingredient of the deep-learning
revival and is still used in many models (the original transformer used
p = 0.1). Forgetting to switch it off at evaluation (model.eval() in
PyTorch) is a classic bug that makes predictions noisy.
Chapter 5
L1 and L2 penalties: a tax on large weights
Add a tax on the size of the weights to the loss, and the model must decide whether each weight earns its keep. L2 (ridge, "weight decay") taxes the square of each weight: big weights pay a lot, small ones almost nothing, so everything shrinks a little but nothing is eliminated. L1 (lasso) charges a flat rate per unit of size: small weights can't justify the fee, so they're wiped out entirely, to exactly zero.
Worked example: with independent (orthonormal) features, the penalized weights have a closed form. Start from the unpenalized weights (3, 0.5, −2) and use penalty strength 1:
| weight | L2: divide by 1 + 1 | L1: move 1 toward zero, stop at zero |
|---|---|---|
| 3 | 1.5 | 2 |
| 0.5 | 0.25 | 0 |
| −2 | −1 | −1 |
Figure 8 · Chart
Fitted weights on ten features where only three matter: L2 leaves small nonzero weights on every useless feature, L1 sets most of them to exactly zero
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| "choose the weights that make the following smallest" | ||
| the features (one row per example) and targets | ||
| the sum of squared prediction errors | ||
| "lambda", the penalty strength | 1 | |
| the L2 norm squared: sum of squared weights | ||
| the L1 norm: sum of absolute weights |
In words: "fit the data, but pay a penalty proportional to the sum of squared weights (L2) or the sum of absolute weights (L1)."
With the numbers: with orthonormal features the solutions are
and
,
the "soft threshold" in soft_threshold.
Level 3: in Python
import math
w = [3, 0.5, -2]
# λ ("lambda" is taken in Python)
lam = 1
# ‖w‖₂², ‖w‖₁
sum(w_i ** 2 for w_i in w), sum(abs(w_i) for w_i in w) # → (13.25, 5.5)
# L2: shrink every weight
[w_i / (1 + lam) for w_i in w] # → [1.5, 0.25, -1.0]
# L1: sign(w) max(|w| - λ, 0)
[math.copysign(max(abs(w_i) - lam, 0), w_i) for w_i in w] # → [2.0, 0.0, -1.0]
In code: penalised_weights gives both closed forms from the table, and
fit_sparse_problem fits the ten-feature problem in the figure (L1 by
alternating gradient steps with soft_threshold).
Why it matters L2 (as weight decay; see AdamW in
primer.ml.optimizers) is on by default when training transformers. L1
is the tool when you want a sparse, interpretable model that uses only a
few features.
Chapter 6
Train, validation and test: practice, mock and real exams
A student has practice questions (to learn from), a mock exam (to check progress and decide what to revise), and the real exam, sat once. Data is split the same way. Training data fits the weights. Validation data tunes choices like model size, learning rate and when to stop. Test data is touched once, at the end, to estimate real-world performance honestly. If you keep tuning against the test set, it quietly becomes a second validation set and stops being honest.
Worked example: 100 examples split 70 / 15 / 15, shuffled first, with no example in two sets. With only 10 examples, 5-fold cross-validation cuts them into 5 folds of 2; each fold takes a turn as the validation set while the model trains on the other 8, so every example is held out exactly once.
Figure 9 · Diagram
flowchart TB
subgraph K["5-fold cross-validation"]
f1["Run 1: [VAL] train train train train"]
f2["Run 2: train [VAL] train train train"]
f3["Run 3: train train [VAL] train train"]
f4["Run 4: train train train [VAL] train"]
f5["Run 5: train train train train [VAL]"]
end
K --> A[Average the 5 validation scores]
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| number of folds | 5 | |
| which fold is held out | 1 to 5 | |
| score(model, fold) | e.g. accuracy of that model on that fold |
In words: "train k models, each with one fold held out, and average their scores on the fold each one didn't see."
With the numbers: 10 examples, : five models, each trained on 8 and scored on 2; if they score 1.0, 0.5, 1.0, 1.0, 0.5, the CV score is 0.8.
Level 3: in Python
examples = list(range(10))
k = 5
# 5 folds of 2
folds = [examples[2 * i: 2 * i + 2] for i in range(k)]
# each model trains on the other 8
[len(examples) - len(fold) for fold in folds] # → [8, 8, 8, 8, 8]
# score of the model trained without fold i, on fold i
scores = [1.0, 0.5, 1.0, 1.0, 0.5]
# (1/k) Σ_i score_i
sum(scores) / k # → 0.8
In code: train_val_test_split shuffles and cuts the indices into three
disjoint sets, and k_fold yields the training and validation indices for
each fold in turn.
Why it matters with small datasets a single split is noisy, and cross-validation gives a more reliable estimate. With large datasets (or expensive models) one fixed validation set is enough.
Chapter 7
Data leakage: the model saw the answers
A student who glimpsed the answer key scores perfectly on the practice exam and learns nothing. Data leakage is when information that won't be available at prediction time, or information from the test set, sneaks into training. Offline scores look superb; production scores collapse.
Worked examples:
- Duplicates across the split. Every row stored twice, labels 30% noisy (so no honest model can beat about 70%). Split before removing duplicates, and a model that just copies its nearest training row scores 87%: it's reading the answers off each test row's twin. Deduplicate first and the same model scores about 50%.
- A feature from the future. A fraud model is given
chargeback_filed, a column filled in only after a customer disputes a charge, which is to say after the outcome. Offline accuracy: 98%. In production that column is always empty at decision time, and accuracy falls to 49%, worse than an honest model using only the legitimate features (61%).
Figure 10 · Diagram
flowchart LR
subgraph T["Timeline of one transaction"]
direction LR
A[Transaction happens] --> P["Model must decide<br/>(prediction time)"]
P --> O[Fraud confirmed?]
O --> C[chargeback_filed recorded]
end
C -. "leaks backwards into<br/>the training table" .-> P
In code: duplicate_leakage_demo and future_feature_demo run both
examples; the "model" in the first is a one-nearest-neighbour lookup (copy the label
of the closest training row), the purest memoriser there is.
Why it matters leakage is the most common reason a model that looked great in evaluation fails in production. For language-model benchmarks the big version is contamination: test questions that appeared in the pretraining data. The defences are procedural: deduplicate before splitting, split by time or by user when the real task is predicting the future or new users, and ask of every feature "would I actually know this at prediction time?"
Test yourself
6 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Training loss drops while validation loss rises. What's happening, and what helps?Think it through, then reveal
Overfitting: the model is memorising training-set noise. Stop early (keep the best-validation weights), add regularization (dropout, weight decay), get more or more varied data, or reduce capacity.
Question 2How do you tell underfitting from overfitting?Think it through, then reveal
Underfitting: training and validation errors are both high. Overfitting: training error is low and validation error is much higher. They need opposite fixes.
Question 3Why does dropout scale the surviving activations by 1/(1 − p)?Think it through, then reveal
So the expected value of each activation is the same during training as at evaluation, where nothing is dropped. The evaluation-time network can then be used as is.
Question 4Why does L1 produce exact zeros but L2 doesn't?Think it through, then reveal
L1's penalty has a constant slope, so small weights feel a fixed pull toward zero that the data can't outweigh; they're thresholded to zero. L2's pull is proportional to the weight, so it fades as a weight shrinks and never quite reaches zero.
Question 5Why must the test set be touched only once?Think it through, then reveal
Every decision made by looking at test scores fits the model to the test set a little. Do it repeatedly and the test score stops being an unbiased estimate of performance on new data.
Question 6Give two examples of data leakage.Think it through, then reveal
The same (or near-duplicate) documents in both the training and test splits; and a feature that's only known after the outcome (a chargeback flag in a fraud model). For LLM evaluation: benchmark questions present in the pretraining data.
Primary sources
The papers behind this lesson
Srivastava, Hinton, Krizhevsky, Sutskever & Salakhutdinov, Dropout: A Simple Way to Prevent Neural Networks from Overfitting (JMLR, 2014): Introduced dropout and showed it as a cheap approximation to averaging many networks.
Read the annotated companion →The paper ↗Tibshirani, Regression Shrinkage and Selection via the Lasso (JRSS B, 1996): Introduced the L1 penalty and its property of setting coefficients exactly to zero.
The paper ↗Geman, Bienenstock & Doursat, Neural Networks and the Bias/Variance Dilemma (Neural Computation, 1992): Framed generalization in neural networks as the bias-variance trade-off.
The paper ↗Belkin, Hsu, Ma & Mandal, Reconciling modern machine learning practice and the bias-variance trade-off (2018): Documented "double descent", where very over-parameterized models generalize better again.
The paper ↗Kapoor & Narayanan, Leakage and the Reproducibility Crisis in ML-based Science (2022): Catalogued the kinds of data leakage and how widely they inflate published results.
The paper ↗Researcher's shelf
Further reading
- Goodfellow, Bengio & Courville, Deep Learning, ch. 7 (regularization): https://www.deeplearningbook.org/contents/regularization.html
- CS231n notes, Neural Networks Part 2 (regularization and dropout): https://cs231n.github.io/neural-networks-2/
- Hastie, Tibshirani & Friedman, The Elements of Statistical Learning (free PDF; ch. 3 and 7): https://hastie.su.domains/ElemStatLearn/
- scikit-learn user guide, Cross-validation: https://scikit-learn.org/stable/modules/cross_validation.html
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.