rumblr Work in progressWIP

● The AI Primer · Lesson 22 · Part 1: how the model works inside

Regularization

learning the pattern instead of memorising the examples

This lesson covers Overfitting, early stopping, dropout, L1/L2, leakage

Members · open during launch 31 min10 figures and diagrams
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Overfitting: training error falls while validation error rises; the model memorised noise. Underfitting: both errors are high.
  2. Fixes for overfitting: more data, early stopping, dropout, weight decay (L2), L1 for sparsity, data augmentation.
  3. Bias is systematic error (too simple); variance is sensitivity to the training sample (too flexible).
  4. Train fits weights, validation tunes choices, test is touched once; cross-validation for small data.
  5. Leakage (duplicates across splits, features from the future) gives great offline scores that collapse in production.

Level 2

How it works, from scratch

Two students prepare for an exam. One learns the ideas. The other memorises last year's paper, answer by answer, and scores 100% on it in practice. On the real exam, with new questions, the first student does fine and the second falls apart. A model that memorises its training examples, noise and all, instead of learning the pattern behind them is overfitting. It looks brilliant on the data it has seen and fails on data it hasn't. Regularization is everything we do to push a model toward the first student.

Worked example: four training points (0, 0), (1, 1), (2, 0), (3, 1), a zig-zag that is really "about 0.5, plus noise".

model error on the 4 training points prediction at x = 4
flat line at 0.5 (the pattern) 0.25 0.5
cubic through all four points (memorised) 0 8

The cubic is perfect on the training data and absurd one step beyond it. You can check the 8 by hand with a difference table: the values 0, 1, 0, 1 have differences 1, −1, 1, then −2, 2, then 4; a cubic keeps that last difference constant, so extending the table gives 6, then 7, then 8.

Figure 1 · Diagram

Reading it: every dataset is a pattern plus noise. A model with too little capacity can't even represent the pattern; one with too much, left unchecked, fits the noise too. Only the middle path does well on new data. Regularization techniques are ways of steering a high-capacity model onto that middle path without giving up its capacity.

Figure 2 · Chart

−1.00 −0.75 −0.50 −0.25 0.00 0.25 0.50 0.75 1.00 x −2.0 −1.5 −1.0 −0.5 0.0 0.5 1.0 1.5 2.0 y Underfit, good fit, memorised true pattern degree 1 degree 3 degree 11 12 training points

Three polynomial fits to 12 noisy sine samples: the degree-1 line misses the curve, the cubic follows it, and the degree-11 polynomial hits every dot but swings wildly between them

Reading it: the dots are 12 noisy samples of the grey sine curve (the true pattern). The straight line (degree 1) can't bend and misses the shape entirely: underfitting. The cubic (degree 3) follows the sine closely. The degree-11 polynomial passes through every dot exactly, and between and beyond them it swings wildly: it has memorised the noise.

The two numbers that tell these apart are errors on two different sets of data:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
how many training and validation examples 4 training points
"for every example in the training set"
the true value 0, 1, 0, 1
the model's prediction 0.5 everywhere for the flat line
mean squared error (see primer.ml.losses) 0.25 for the flat line

In words: "training error is the average squared miss on the data the model learned from; validation error is the same average on data it has never seen."

With the numbers: flat line: . Cubic: every miss is 0, so training error is 0; its validation error on any point beyond x = 3 is enormous.

Level 3: in Python
# the true values at x = 0, 1, 2, 3
y = [0, 1, 0, 1]
# the flat line predicts 0.5 everywhere
y_hat = [0.5, 0.5, 0.5, 0.5]
n = len(y)
# (1/n) Σ (y_i - ŷ_i)²
sum((y_i - y_hat_i) ** 2 for y_i, y_hat_i in zip(y, y_hat)) / n  # → 0.25

In code: zigzag_fit_error and zigzag_predict fit a polynomial of any degree to the four zig-zag points and report its training error and its prediction; true_function and make_curve_data draw the noisy sine samples in the figure.

Why it matters training error alone always rewards memorisation. The gap between training and validation error is the single most useful diagnostic in machine learning.

Chapter 1

Underfitting vs. overfitting: finding the middle

Goldilocks tries three bowls of porridge: too cold, too hot, just right. Model capacity works the same way, and you find "just right" by measuring, not guessing: sweep the capacity and watch both errors.

Worked example: 12 noisy points from a sine wave, 300 fresh points for validation, polynomials of increasing degree.

degree training error validation error diagnosis
1 0.20 0.30 underfit: both high
3 0.021 0.065 about right
5 0.013 0.057 best: lowest validation error
9 0.004 0.71 overfit: validation 12× worse than at degree 5
11 0.0000 7,631 memorised

Figure 3 · Chart

0 2 4 6 8 10 polynomial degree (model capacity) 1 0 − 5 1 0 − 3 1 0 − 1 1 0 1 1 0 3 mean squared error (log) Training error only falls; validation error is U-shaped training error validation error

Error against polynomial degree on a log scale: training error only falls, while validation error is lowest at degree 5 and then climbs steeply

Reading it: the horizontal axis is model capacity (polynomial degree); the vertical axis is error on a log scale. The training curve (blue) only ever goes down: more capacity always fits the training data better. The validation curve (orange) falls, is nearly flat from degree 3 to 5 (lowest at 5), then climbs from degree 6 onward and shoots up. The left side of that valley is underfitting, the right side overfitting; the bottom is the model you want.

In code: polynomial_errors fits by least squares (choosing the coefficients that minimise training MSE) and reports both errors.

Why it matters "both errors high" and "training low, validation high" need opposite fixes. Underfitting wants more capacity, more features or more training; overfitting wants more data, regularization or early stopping. Diagnosing which one you have is the point.

Chapter 2

Early stopping: take the cake out before it burns

A baker checks the cake with a toothpick every few minutes and takes it out when it comes out clean. Leave it longer and it burns, however good it was a minute ago. Training a flexible model is similar: validation error falls while the model learns the pattern, then rises once it starts memorising. Early stopping watches validation error during training, stops once it has failed to improve for a few checks (the patience), and keeps the weights from the best check.

Worked example: validation losses per epoch 1.0, 0.8, 0.6, 0.55, 0.58, 0.65, 0.7 with patience 2. Epoch 3 is the best. Epochs 4 and 5 are both worse, so training stops at epoch 5 and the weights from epoch 3 are kept.

Figure 4 · Diagram

Reading it: after every epoch the loop asks one question: is this the best validation loss so far? If yes, snapshot the weights. If not, count a strike. Reaching the patience limit ends training and restores the snapshot, so you always ship the best model seen, not the last one.

Figure 5 · Chart

0 250 500 750 1000 1250 1500 1750 2000 epoch 0.2 0.4 0.6 0.8 cross-entropy An over-sized network on 30 noisy points best validation (epoch 314) training loss validation loss

Loss curves for an over-sized network: training loss keeps falling while validation loss bottoms out near epoch 300 and then rises

Reading it: a 64-unit network trained on only 30 noisy points. Both curves fall at first. Around epoch 300 (the dashed line) the validation curve bottoms out and turns upward while the training curve keeps falling: the network has started fitting individual noisy points. Everything to the right of the dashed line is wasted, or worse. Early stopping keeps the weights at the dashed line.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the epoch number 0 to 6
validation loss after epoch 0.55 at
"the at which the following is smallest" (the position, not the value) 3
"t-star", the best epoch so far 3
patience how many non-improving epochs to tolerate 2

In words: "remember the epoch with the lowest validation loss, and stop once you've gone patience epochs past it without doing better."

With the numbers: ; at , , so stop.

Level 3: in Python
# validation loss after epoch t = 0, 1, 2, ...
L_val = [1.0, 0.8, 0.6, 0.55, 0.58, 0.65, 0.7]
patience = 2
t_star = 0
for t, loss in enumerate(L_val):
    if loss < L_val[t_star]:
        # a new best epoch: remember it
        t_star = t
    if t - t_star >= patience:
        # patience used up: stop here
        break
t_star, t  # → (3, 5)

In code: early_stopping replays a run's validation losses and returns the best epoch and the epoch where training stops; train_flexible_model trains the over-sized network in the figure and records both losses.

Why it matters it's the cheapest regularizer there is: no change to the model, just a rule for when to stop. It's used almost everywhere a model is trained for multiple epochs, including fine-tuning language models.

Chapter 3

Bias and variance: two ways to miss the target

Picture two archers. The first groups every arrow tightly, but a hand's width to the left of the bullseye: consistently wrong. That's bias. The second's arrows are centred on the bullseye on average but scattered all over the target: inconsistent. That's variance. A model trained on a different sample of data is like another volley of arrows. Simple models behave like the first archer; very flexible ones like the second.

Worked example: fit polynomials to 300 different random training sets of 30 points each, and measure (averaged over the input range):

degree bias² (systematic error) variance (swing between training sets)
1 0.174 0.022
3 0.0035 0.015
9 0.0038 4.5

Figure 6 · Chart

−1.00 −0.75 −0.50 −0.25 0.00 0.25 0.50 0.75 1.00 x −2.0 −1.5 −1.0 −0.5 0.0 0.5 1.0 1.5 2.0 prediction degree 1: bias² 0.174, variance 0.022 −1.00 −0.75 −0.50 −0.25 0.00 0.25 0.50 0.75 1.00 x degree 9: bias² 0.004, variance 4.545

Twenty fits from different training sets: degree-1 lines agree but all miss the sine (high bias), degree-9 curves average to the sine but scatter widely (high variance)

Reading it: each thin line is the model fitted to one random training set; the thick grey curve is the truth. On the left (degree 1), the lines agree with each other but all miss the sine's shape in the same way: high bias, low variance. On the right (degree 9), the lines follow the sine on average, but each one wiggles differently, especially near the edges: low bias, high variance.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the true pattern
"f-hat", the model's prediction, which depends on which training set it saw one thin line
expected value: the average over many random training sets (and noise) average of the 300 fits
a noisy observation, plus noise
the noise variance: error no model can remove

In words: "a model's expected squared error on new data splits into how far its average prediction is from the truth (bias squared), plus how much its predictions scatter around their own average (variance), plus noise nobody can predict."

With the numbers: degree 1: ; degree 3: ; degree 9: . Degree 3 has the best total.

Level 3: in Python
sigma = 0.3
# σ²: the error no model can remove
noise = sigma ** 2
for degree, bias_sq, variance in [(1, 0.174, 0.022), (3, 0.0035, 0.015), (9, 0.0038, 4.5)]:
    # bias² + variance + σ²
    total = bias_sq + variance + noise
    # two significant figures
    print(degree, f"{total:.2g}")  # → 1 0.29 3 0.11 9 4.6

In code: bias_variance fits one polynomial to each of many random training sets and splits their error into bias², variance and noise.

Why it matters it names the two failure modes and explains why more data helps flexible models (averaging tames variance) but not rigid ones (more data doesn't fix bias). Very large neural networks complicate the picture ("double descent": past a point, even bigger models generalize better again), but the vocabulary is universal.

Chapter 4

Dropout: nobody gets to be indispensable

A coach who randomly sends half the team home before each practice forces every player to learn every position; no single star can carry the team. Dropout does this to neurons: during training, each activation is set to zero with probability p on every step, so the network can't rely on any single neuron or fragile combination of neurons.

Worked example: four activations (1, 1, 1, 1), drop rate p = 0.5. Suppose the random mask keeps the last two: (0, 0, 1, 1). The survivors are scaled by 1 / (1 − 0.5) = 2, giving (0, 0, 2, 2). The average is still 1, so the next layer sees the same size signal on average. At evaluation time, nothing is dropped and nothing is scaled.

Figure 7 · Diagram

Reading it: on the left, one training step: two of the four units were dropped to 0 and the two survivors were doubled, so the total (4) matches what the full layer would have sent. A different random pair is dropped on the next step. On the right, at evaluation time every unit passes through unchanged. Scaling during training ("inverted dropout") is what lets evaluation be a plain pass-through.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
a layer's activations (1, 1, 1, 1)
"h-tilde", the activations after dropout (0, 0, 2, 2)
the random mask of 0s and 1s (0, 0, 1, 1)
each mask entry is independently 1 with probability , else 0 (a coin flip) 50/50
the drop rate 0.5
multiply entry by entry

In words: "flip a biased coin for each activation, zero the ones that lose, and divide the survivors by the keep probability."

With the numbers: .

Level 3: in Python
import random
random.seed(0)
h = [1, 1, 1, 1]
p = 0.5
# m_i ~ Bernoulli(1 - p): a coin flip each
m = [1 if random.random() < 1 - p else 0 for _ in h]
m  # → [0, 0, 1, 1]
# (m ⊙ h) / (1 - p)
[m_i * h_i / (1 - p) for m_i, h_i in zip(m, h)]  # → [0.0, 0.0, 2.0, 2.0]

In code: dropout draws the mask and scales the survivors during training, and passes activations through unchanged at evaluation time.

Why it matters dropout was a key ingredient of the deep-learning revival and is still used in many models (the original transformer used p = 0.1). Forgetting to switch it off at evaluation (model.eval() in PyTorch) is a classic bug that makes predictions noisy.

Chapter 5

L1 and L2 penalties: a tax on large weights

Add a tax on the size of the weights to the loss, and the model must decide whether each weight earns its keep. L2 (ridge, "weight decay") taxes the square of each weight: big weights pay a lot, small ones almost nothing, so everything shrinks a little but nothing is eliminated. L1 (lasso) charges a flat rate per unit of size: small weights can't justify the fee, so they're wiped out entirely, to exactly zero.

Worked example: with independent (orthonormal) features, the penalized weights have a closed form. Start from the unpenalized weights (3, 0.5, −2) and use penalty strength 1:

weight L2: divide by 1 + 1 L1: move 1 toward zero, stop at zero
3 1.5 2
0.5 0.25 0
−2 −1 −1

Figure 8 · Chart

0 1 2 3 4 5 6 7 8 9 feature (only 0, 1, 2 matter) −2 −1 0 1 2 3 fitted weight L1 zeroes useless weights; L2 only shrinks them true weights L2 (ridge) L1 (lasso)

Fitted weights on ten features where only three matter: L2 leaves small nonzero weights on every useless feature, L1 sets most of them to exactly zero

Reading it: ten features, but only the first three affect the target (grey bars show the true weights: 3, −2, 1.5, then zeros). The L2 fit (blue) shrinks everything a little and leaves small nonzero weights on all seven useless features. The L1 fit (orange) gets the three real weights nearly right and sets most of the useless ones to exactly zero: it has selected features for you.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
"choose the weights that make the following smallest"
the features (one row per example) and targets
the sum of squared prediction errors
"lambda", the penalty strength 1
the L2 norm squared: sum of squared weights
the L1 norm: sum of absolute weights

In words: "fit the data, but pay a penalty proportional to the sum of squared weights (L2) or the sum of absolute weights (L1)."

With the numbers: with orthonormal features the solutions are and , the "soft threshold" in soft_threshold.

Level 3: in Python
import math
w = [3, 0.5, -2]
# λ ("lambda" is taken in Python)
lam = 1
# ‖w‖₂², ‖w‖₁
sum(w_i ** 2 for w_i in w), sum(abs(w_i) for w_i in w)  # → (13.25, 5.5)
# L2: shrink every weight
[w_i / (1 + lam) for w_i in w]  # → [1.5, 0.25, -1.0]
# L1: sign(w) max(|w| - λ, 0)
[math.copysign(max(abs(w_i) - lam, 0), w_i) for w_i in w]  # → [2.0, 0.0, -1.0]

In code: penalised_weights gives both closed forms from the table, and fit_sparse_problem fits the ten-feature problem in the figure (L1 by alternating gradient steps with soft_threshold).

Why it matters L2 (as weight decay; see AdamW in primer.ml.optimizers) is on by default when training transformers. L1 is the tool when you want a sparse, interpretable model that uses only a few features.

Chapter 6

Train, validation and test: practice, mock and real exams

A student has practice questions (to learn from), a mock exam (to check progress and decide what to revise), and the real exam, sat once. Data is split the same way. Training data fits the weights. Validation data tunes choices like model size, learning rate and when to stop. Test data is touched once, at the end, to estimate real-world performance honestly. If you keep tuning against the test set, it quietly becomes a second validation set and stops being honest.

Worked example: 100 examples split 70 / 15 / 15, shuffled first, with no example in two sets. With only 10 examples, 5-fold cross-validation cuts them into 5 folds of 2; each fold takes a turn as the validation set while the model trains on the other 8, so every example is held out exactly once.

Figure 9 · Diagram

Reading it: the data is cut into five equal folds. In each run, one fold is held out (VAL) and the model is trained from scratch on the other four. Every example is validated on exactly once, and the average of the five scores is a steadier estimate than any single split.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
number of folds 5
which fold is held out 1 to 5
score(model, fold) e.g. accuracy of that model on that fold

In words: "train k models, each with one fold held out, and average their scores on the fold each one didn't see."

With the numbers: 10 examples, : five models, each trained on 8 and scored on 2; if they score 1.0, 0.5, 1.0, 1.0, 0.5, the CV score is 0.8.

Level 3: in Python
examples = list(range(10))
k = 5
# 5 folds of 2
folds = [examples[2 * i: 2 * i + 2] for i in range(k)]
# each model trains on the other 8
[len(examples) - len(fold) for fold in folds]  # → [8, 8, 8, 8, 8]
# score of the model trained without fold i, on fold i
scores = [1.0, 0.5, 1.0, 1.0, 0.5]
# (1/k) Σ_i score_i
sum(scores) / k  # → 0.8

In code: train_val_test_split shuffles and cuts the indices into three disjoint sets, and k_fold yields the training and validation indices for each fold in turn.

Why it matters with small datasets a single split is noisy, and cross-validation gives a more reliable estimate. With large datasets (or expensive models) one fixed validation set is enough.

Chapter 7

Data leakage: the model saw the answers

A student who glimpsed the answer key scores perfectly on the practice exam and learns nothing. Data leakage is when information that won't be available at prediction time, or information from the test set, sneaks into training. Offline scores look superb; production scores collapse.

Worked examples:

  • Duplicates across the split. Every row stored twice, labels 30% noisy (so no honest model can beat about 70%). Split before removing duplicates, and a model that just copies its nearest training row scores 87%: it's reading the answers off each test row's twin. Deduplicate first and the same model scores about 50%.
  • A feature from the future. A fraud model is given chargeback_filed, a column filled in only after a customer disputes a charge, which is to say after the outcome. Offline accuracy: 98%. In production that column is always empty at decision time, and accuracy falls to 49%, worse than an honest model using only the legitimate features (61%).

Figure 10 · Diagram

Reading it: read the timeline left to right. The model has to decide at the second box. The chargeback is only recorded at the last box, after the outcome is known. In a historical training table every column sits side by side, so nothing stops the model from using a column that, in real time, it could never have had. The dotted arrow is the leak.

In code: duplicate_leakage_demo and future_feature_demo run both examples; the "model" in the first is a one-nearest-neighbour lookup (copy the label of the closest training row), the purest memoriser there is.

Why it matters leakage is the most common reason a model that looked great in evaluation fails in production. For language-model benchmarks the big version is contamination: test questions that appeared in the pretraining data. The defences are procedural: deduplicate before splitting, split by time or by user when the real task is predicting the future or new users, and ask of every feature "would I actually know this at prediction time?"

Test yourself

6 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Training loss drops while validation loss rises. What's happening, and what helps?Think it through, then reveal

Overfitting: the model is memorising training-set noise. Stop early (keep the best-validation weights), add regularization (dropout, weight decay), get more or more varied data, or reduce capacity.

Question 2How do you tell underfitting from overfitting?Think it through, then reveal

Underfitting: training and validation errors are both high. Overfitting: training error is low and validation error is much higher. They need opposite fixes.

Question 3Why does dropout scale the surviving activations by 1/(1 − p)?Think it through, then reveal

So the expected value of each activation is the same during training as at evaluation, where nothing is dropped. The evaluation-time network can then be used as is.

Question 4Why does L1 produce exact zeros but L2 doesn't?Think it through, then reveal

L1's penalty has a constant slope, so small weights feel a fixed pull toward zero that the data can't outweigh; they're thresholded to zero. L2's pull is proportional to the weight, so it fades as a weight shrinks and never quite reaches zero.

Question 5Why must the test set be touched only once?Think it through, then reveal

Every decision made by looking at test scores fits the model to the test set a little. Do it repeatedly and the test score stops being an unbiased estimate of performance on new data.

Question 6Give two examples of data leakage.Think it through, then reveal

The same (or near-duplicate) documents in both the training and test splits; and a feature that's only known after the outcome (a chargeback flag in a fraud model). For LLM evaluation: benchmark questions present in the pretraining data.

Primary sources

The papers behind this lesson

Srivastava, Hinton, Krizhevsky, Sutskever & Salakhutdinov, Dropout: A Simple Way to Prevent Neural Networks from Overfitting (JMLR, 2014): Introduced dropout and showed it as a cheap approximation to averaging many networks.

Read the annotated companion →The paper ↗

Tibshirani, Regression Shrinkage and Selection via the Lasso (JRSS B, 1996): Introduced the L1 penalty and its property of setting coefficients exactly to zero.

The paper ↗

Geman, Bienenstock & Doursat, Neural Networks and the Bias/Variance Dilemma (Neural Computation, 1992): Framed generalization in neural networks as the bias-variance trade-off.

The paper ↗

Belkin, Hsu, Ma & Mandal, Reconciling modern machine learning practice and the bias-variance trade-off (2018): Documented "double descent", where very over-parameterized models generalize better again.

The paper ↗

Kapoor & Narayanan, Leakage and the Reproducibility Crisis in ML-based Science (2022): Catalogued the kinds of data leakage and how widely they inflate published results.

The paper ↗

Researcher's shelf

Further reading

  • Goodfellow, Bengio & Courville, Deep Learning, ch. 7 (regularization): https://www.deeplearningbook.org/contents/regularization.html
  • CS231n notes, Neural Networks Part 2 (regularization and dropout): https://cs231n.github.io/neural-networks-2/
  • Hastie, Tibshirani & Friedman, The Elements of Statistical Learning (free PDF; ch. 3 and 7): https://hastie.su.domains/ElemStatLearn/
  • scikit-learn user guide, Cross-validation: https://scikit-learn.org/stable/modules/cross_validation.html

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.