rumblr Work in progressWIP

● The AI Primer · Lesson 9 · Part 1: how the model works inside

Training stages

how a text predictor becomes an assistant

This lesson covers Pretraining, SFT, RLHF and DPO, LoRA, fine-tuning vs. RAG

Members · open during launch 35 min13 figures and diagrams8 interactive
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Pretraining predicts the next token over huge text and gives knowledge; SFT teaches the assistant format; preference tuning shapes tone, helpfulness and safety.
  2. SFT is next-token loss with the prompt masked out.
  3. RLHF trains a reward model on pairwise preferences, then optimises the LM against it; DPO gets the same effect directly from preference pairs.
  4. LoRA trains a tiny low-rank correction B·A beside frozen weights: 0.39% of a layer's parameters at rank 8, 3.1% at rank 64, mergeable for zero added latency.
  5. Fine-tuning teaches behaviour; RAG supplies knowledge. Start with prompting and RAG.
  6. Distillation trains a small student on a big teacher's soft targets; it's a major cost lever.

Level 2

How it works, from scratch

A chat assistant is built in layers, and each layer explains a different behaviour you see:

Figure 1 · Diagram

Reading it: read left to right; each box is a training stage and each arrow hands a model to the next stage. The model vendor runs the first three stages: pretraining (months, millions of dollars) gives a base model that knows facts but just continues text; supervised fine-tuning (SFT) teaches it the format of being an assistant; preference tuning shapes tone, helpfulness and safety. You work in the last box. When a model knows a fact it learned that in pretraining; when it answers in tidy bullet points it learned that in SFT and preference tuning.

This lesson builds a small, real version of every stage: the pretraining and SFT losses, a reward model, DPO, LoRA, the "which adaptation?" decision, and distillation.

The pipeline above, with a question you can ask it: which stage taught the assistant that?

Chapter 1

Pretraining: guess the next word, a trillion times

Make the guesses yourself first; the loss will then read as what it is, an average penalty.

Everyday picture Imagine reading every book in a library with a card covering the next word, guessing it, then sliding the card to check. Do that trillions of times and you absorb grammar, facts, arithmetic habits and coding conventions, because every one of them helps you guess the next word.

Tiny worked example A 5-token sequence gives 4 guesses (token 1 predicts token 2, and so on; the last token has nothing after it to check). Suppose the model gave the right next token probability 0.25, 0.25, 0.5 and 0.5. The penalty for each guess is −ln(probability): 1.386, 1.386, 0.693, 0.693. The pretraining loss is their average, 1.040.

Figure 2 · Diagram

Reading it: every position makes one guess about the token that follows it, using only what came before. The model is scored on all guesses at once and the average penalty is the number training pushes down. One text sequence of n tokens therefore gives n − 1 training examples for free, which is why unlabeled text is such a rich training signal.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
the loss: one number, lower is better ≥ 0
number of tokens in the sequence integer
add up the term for every position i from 1 to n−1
the i-th token (an integer id) 0 … vocab−1
all tokens up to and including position i (the context)
probability the model with weights θ gives the true next token, given the context. The bar "∣" reads "given" 0 … 1
all the model's weights millions to trillions of numbers
natural logarithm. ln 1 = 0 and ln of a small number is very negative, so −log turns "probability of being right" into "penalty"

In words: the loss is the average, over every position, of minus the log of the probability the model gave to the token that actually came next.

On the worked example: n = 5, the four probabilities are 0.25, 0.25, 0.5, 0.5, so the loss is −(ln 0.25 + ln 0.25 + ln 0.5 + ln 0.5) / 4 = (1.386 + 1.386 + 0.693 + 0.693) / 4 = 1.040.

Level 3: in Python
import math
# p_θ(t_{i+1} | t_{1..i}) for each of the guesses
p = [0.25, 0.25, 0.5, 0.5]
# 5 tokens give n − 1 = 4 guesses
n = len(p) + 1
[round(-math.log(p_i), 3) for p_i in p]  # → [1.386, 1.386, 0.693, 0.693]
L_pretrain = -sum(math.log(p_i) for p_i in p) / (n - 1)
print(f"{L_pretrain:.3f}")  # → 1.040

In code: per_position_losses computes −ln p for every next-token guess (with a stable log_softmax), and next_token_loss averages them into the pretraining loss.

Why it matters in practice. Pretraining is where knowledge comes from, and it is frozen at a date (the "knowledge cutoff"). A base model continues text instead of answering questions: ask "What is the capital of France?" and it may continue with "What is the capital of Spain?", because question lists are common on the web.

Chapter 2

Supervised fine-tuning (SFT): study the answers, not the questions

Everyday picture An apprentice studies a binder of worked examples: customer question, then the expert's reply. They read the question carefully, but they are graded only on writing the reply. Nobody marks them on reproducing the customer's typing.

Tiny worked example Same 5 tokens, but the first 3 are the prompt and the last 2 are the reply. Only the 2 guesses whose target is a reply token count: penalties 0.693 and 0.693, average 0.693. The prompt guesses (1.386 each) are ignored.

Figure 3 · Diagram

Reading it: the whole sequence flows through the model, so the reply is conditioned on the prompt. The dotted arrows show which predictions reach the loss: only those that produce reply tokens. That mask (a list of true/false per position) is the whole difference between pretraining code and SFT code.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
the mask: 1 if token i+1 belongs to the reply, else 0 0 or 1
how many predictions are graded (the reply length) integer
everything else as in the pretraining formula

In words: average the next-token penalty over the reply tokens only.

On the worked example: m = (0, 0, 1, 1), so the loss is (0.693 + 0.693) / 2 = 0.693.

Level 3: in Python
import math
p = [0.25, 0.25, 0.5, 0.5]
# 1 only where the target is a reply token
m = [0, 0, 1, 1]
L_SFT = -sum(m_i * math.log(p_i) for m_i, p_i in zip(m, p)) / sum(m)
round(L_SFT, 3)  # → 0.693

Figure 4 · Chart

t1→t2 t2→t3 t3→t4 t4→t5 prediction (grey = prompt, masked; blue = reply) 0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 1.6 1.8 penalty −ln p(true next token) SFT grades only the reply pretraining loss 1.040 (all positions) SFT loss 0.693 (reply only)

Only the two reply positions count: SFT loss 0.693 versus 1.040 averaged over every position

Reading it: each bar is one next-token prediction from the worked example, and its height is the penalty −ln p. Grey bars predict prompt tokens and are masked out of the SFT loss; blue bars predict reply tokens and are the only ones that count. The dashed lines mark the two averages: 1.040 over everything (pretraining) and 0.693 over the reply (SFT).

In code: response_mask builds the true/false list of graded predictions, and sft_loss averages per_position_losses over only the positions it marks.

Why it matters in practice. SFT needs far less data than pretraining (thousands to hundreds of thousands of examples) and quality beats quantity: a few thousand excellent examples outperform a mountain of mediocre ones. If you fine-tune your own model and forget the mask, it learns to write user prompts too.

The mask above, with the boundary in your hands.

Chapter 3

Preference tuning: a taste test instead of a recipe

Score two answers yourself first: the sigmoid will then read as the taste test it is.

Everyday picture It is hard to write down the perfect answer, but easy to taste two dishes and say which is better. Preference tuning collects exactly those judgements: people (or a model following written principles) compare two responses and pick one.

3a. RLHF: train a critic, then train the cook against it

A reward model learns to predict those judgements, giving each response a score. Reinforcement learning then tunes the language model to earn high scores. The link between scores and "which one wins" is the Bradley-Terry model, built on the sigmoid function σ(z) = 1 / (1 + e^−z), which squashes any number into a probability between 0 and 1 (σ(0) = 0.5, σ(2) = 0.881).

Tiny worked example The reward model scores the chosen answer 3.0 and the rejected one 1.0. The gap is 2, so it predicts the chosen answer wins with probability σ(2) = 0.881, and its loss on this pair is −ln 0.881 = 0.127. Had it scored both 1.0, it would predict a coin flip (0.5) and pay ln 2 = 0.693.

Figure 5 · Diagram

Reading it: the loop has two learners. First, labelers compare pairs and the reward model learns to agree with them. Second, the language model generates new responses, the reward model scores them, and reinforcement learning (usually PPO) nudges the language model towards higher scores, with a penalty for drifting far from where it started. Two models, two training runs, lots of moving parts.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
the preferred ("winning") and rejected ("losing") responses text
"is preferred to"
the reward model's score for response y any real number
sigmoid, : turns a score gap into a probability 0 … 1
Euler's number, 2.718…; is "e to the power −z"
the reward model's loss on this pair ≥ 0

In words: the chance the preferred answer wins is the sigmoid of the score gap, and the reward model is penalised by minus the log of the probability it gave to the choice people actually made.

On the worked example: r(y_w) = 3, r(y_l) = 1, gap 2, σ(2) = 0.881, loss −ln 0.881 = 0.127.

Level 3: in Python
import math
def sigma(z):
    return 1 / (1 + math.exp(-z))
r_w, r_l = 3.0, 1.0
# P(y_w ≻ y_l)
round(sigma(r_w - r_l), 3)  # → 0.881
# L_RM
round(-math.log(sigma(r_w - r_l)), 3)  # → 0.127
# equal scores: a coin flip, ln 2
round(-math.log(sigma(1.0 - 1.0)), 3)  # → 0.693

In code: sigmoid squashes a score gap into a probability, preference_probability applies it to two rewards (the Bradley-Terry model), and reward_model_loss is minus the log of that probability.

3b. DPO: skip the critic

Everyday picture Instead of hiring a food critic and cooking to please them, edit the recipe book directly from the taste-test results, while keeping a copy of the original book so you don't drift too far from it.

Tiny worked example A log-probability is the log of the probability the model gives a whole response (a sum of per-token log probabilities; more negative means less likely). The frozen reference model gives both answers −11. After some training the policy gives the chosen answer −10 (more likely than before) and the rejected one −12 (less likely). With β = 0.1 the margin is 0.1 × ((−10 − −11) − (−12 − −11)) = 0.1 × (1 + 1) = 0.2; the loss is −ln σ(0.2) = 0.598, down from ln 2 = 0.693 when the policy still equalled the reference.

Figure 6 · Diagram

Reading it: each preference pair is scored twice: by the model being trained and by a frozen copy of where it started. The margin measures how much more the policy has boosted the chosen answer than the rejected one, relative to the reference. The loss is the reward-model loss again, but the "reward" is read straight off the policy's own probabilities, so there is no separate reward model and no reinforcement-learning loop.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
the policy: the model being trained, with weights θ; π(y) is the probability it gives response y 0 … 1
the frozen reference model (usually the SFT model) 0 … 1
the log-ratio: how much more likely training has made y. β times it is the implicit reward , so the bracket scaled by β is any real
beta, the leash to the reference: it sets how much a change in log-probability counts, so a larger β satisfies the loss with a smaller departure and keeps the policy closer to the reference; a smaller β lets the preferences pull it further away typically 0.1 to 0.5
, sigmoid and natural log, as above

In words: raise the probability of the chosen answer and lower the rejected one, measured relative to the frozen starting model, and penalise the model by minus the log-sigmoid of that scaled gap.

On the worked example: β = 0.1, log-ratios +1 (chosen) and −1 (rejected), so implicit rewards r̂ = 0.1 × 1 = +0.1 and 0.1 × (−1) = −0.1, margin 0.1 − (−0.1) = 0.2, σ(0.2) = 0.550, loss 0.598. The gradient's size is β × (1 − σ(margin)) = 0.1 × 0.450 = 0.045: pairs the policy already ranks correctly get gentle updates, and pairs it ranks the wrong way get strong ones.

Level 3: in Python
import math
def sigma(z):
    return 1 / (1 + math.exp(-z))
beta = 0.1
# chosen answer: policy, reference
logpi_w, logpi_ref_w = -10.0, -11.0
# rejected answer: policy, reference
logpi_l, logpi_ref_l = -12.0, -11.0
# implicit rewards r̂ = β × log-ratio
r_w = beta * (logpi_w - logpi_ref_w)
r_l = beta * (logpi_l - logpi_ref_l)
print(f"{r_w:.1f} {r_l:.1f}")  # → 0.1 -0.1
margin = r_w - r_l
print(f"{margin:.1f} {sigma(margin):.3f}")  # → 0.2 0.550
# L_DPO
round(-math.log(sigma(margin)), 3)  # → 0.598
# how hard this pair pushes
round(beta * (1 - sigma(margin)), 3)  # → 0.045

Figure 7 · Chart

−6 −4 −2 0 2 4 6 DPO margin (how strongly the policy already prefers the chosen answer) 0.00 0.02 0.04 0.06 0.08 0.10 update strength β(1 − σ(margin)) DPO pushes hardest on pairs it gets wrong (β = 0.1) worked example: margin 0.2 → 0.045

DPO's push is β for pairs ranked backwards, half that at margin 0, and fades to zero once a pair is learned

Reading it: the x-axis is the DPO margin (how strongly the policy already prefers the chosen answer, relative to the reference); the y-axis is how hard one gradient step pushes. Far left, the policy has the pair backwards and gets the full push, β. At zero (no preference yet) it gets half. Far right, the pair is learned and updates fade to nothing, so training effort flows automatically to the pairs still wrong.

Figure 8 · Chart

0 25 50 75 100 125 150 175 200 training step 0.0 0.2 0.4 0.6 0.8 1.0 policy probability DPO on three candidate answers helpful and correct rude correct but rambling

The helpful answer climbs towards 1 while rude falls fastest and rambling falls more slowly, staying above rude

Reading it: a one-prompt toy policy starts uniform over three answers (1/3 each). The preference data says "helpful and correct" beats both others and "correct but rambling" beats "rude". As training steps pass (x-axis), probability (y-axis) flows to the top answer, the rude answer is pushed down fastest, and the rambling answer falls too but more slowly (0.166 at step 50, 0.005 by step 200), staying above rude the whole way: the ordering people expressed. Nothing here stops the top answer taking nearly everything, because every pair it appears in keeps pushing it up.

In code: dpo_margin computes the gap between implicit rewards (β times each log-ratio), dpo_loss turns it into −log σ(margin), and dpo_update_strength gives the push β × (1 − σ(margin)) plotted in the first figure; train_toy_dpo trains the three-answer toy policy of the second.

Why it matters in practice. DPO is simpler and more stable than RLHF, which is why it is widely used in open-model fine-tuning. Constitutional AI and AI-feedback methods scale the labelling by having a model judge responses against written principles. This stage is where tone, helpfulness, refusals and safety behaviour mostly come from.

DPO's worked example and its toy policy, with the leash in your hands.

Chapter 4

LoRA: sticky notes instead of reprinting the textbook

Do the worked example with your own numbers first; the two paths will then read as what they are.

Everyday picture You want to adapt a 1,000-page textbook for your class. Reprinting it is expensive. Instead you add a small stack of sticky notes with corrections. The book stays untouched; the notes are cheap to write, store and swap, and you can photocopy them into the book when you're done.

Tiny worked example Frozen weights W = the 2×2 identity (it copies its input). Adapter B = (1, 0) as a column and A = (0, 1) as a row. For input x = (1, 2): the frozen path gives W·x = (1, 2). The adapter first squeezes x to one number, A·x = 2, then expands it back, B·2 = (2, 0). The output is (1, 2) + (2, 0) = (3, 2). Two small vectors changed the layer's behaviour without touching W.

Figure 9 · Diagram

Reading it: the input takes two paths. The top path is the pretrained layer, which never changes. The bottom path is the adapter: A squeezes the d-dimensional input down to a tiny rank r (say 8), and B expands it back. Because B starts at zero, the bottom path adds nothing at first and the model starts exactly where pretraining left it. Only A and B are trained. Afterwards you can add B·A into W once ("merge") and serve the model at exactly the original speed.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
input row vector (d_in,)
frozen pretrained weight matrix (d_out, d_in)
transpose: flip rows and columns so the shapes line up for multiplication
trainable "down" matrix, random at start (r, d_in)
trainable "up" matrix, zero at start (d_out, r)
the rank: the width of the bottleneck usually 4 to 64
alpha, a scale knob; α/r keeps update size steady when you change r often r or 2r

In words: the output is what the frozen layer produces plus a scaled correction that passes through a narrow r-number bottleneck.

On the worked example: W = I, A = (0, 1), B = (1, 0)ᵀ, α/r = 1, x = (1, 2): xWᵀ = (1, 2), xAᵀ = 2, 2·Bᵀ = (2, 0), y = (3, 2).

Level 3: in Python
# x Mᵀ: the dot product of x with each row of M
def times_transpose(x, M):
    return [sum(x_k * m_k for x_k, m_k in zip(x, row)) for row in M]
# frozen, d_out × d_in
W = [[1, 0], [0, 1]]
# r × d_in, with r = 1
A = [[0, 1]]
# d_out × r
B = [[1], [0]]
x, alpha_over_r = [1, 2], 1
# x Wᵀ
frozen = times_transpose(x, W)
# x Aᵀ: squeezed to r numbers
squeezed = times_transpose(x, A)
# (x Aᵀ) Bᵀ: expanded back
correction = times_transpose(squeezed, B)
frozen, squeezed, correction  # → ([1, 2], [2], [2, 0])
# y
[f_j + alpha_over_r * c_j for f_j, c_j in zip(frozen, correction)]  # → [3, 2]
d_in = d_out = 4096
# full, then LoRA at each rank
d_out * d_in, [r * d_in + d_out * r for r in (64, 16, 8)]  # → (16777216, [524288, 131072, 65536])

Parameter savings for one 4096 × 4096 layer:

Method Trainable parameters Share of full
full fine-tune 16,777,216 100%
LoRA r = 64 524,288 3.1%
LoRA r = 16 131,072 0.78%
LoRA r = 8 65,536 0.39%

Figure 10 · Chart

0 50 100 150 200 250 300 training step 1 0 − 8 1 0 − 6 1 0 − 4 1 0 − 2 1 0 0 mean squared error (log scale) The task needs a rank-2 change: rank 1 can't express it rank 1 rank 2 rank 4

A rank-1 adapter plateaus while ranks 2 and 4 drive the error to essentially zero

Reading it: the task needs a rank-2 change to a frozen 16 × 16 layer. The y-axis (log scale) is the training error; the x-axis is the training step. A rank-1 adapter plateaus: its bottleneck is too narrow to express the change. Ranks 2 and 4 drive the error to essentially zero. This is the LoRA bet: the change a fine-tune needs is low-rank even though the weights themselves are not.

In code: LoRALinear holds the frozen W beside the trainable A and B, and LoRALinear.merged_weight folds the adapter into W for serving. lora_trainable_params and full_trainable_params count the table's parameters, and train_toy_lora trains the adapters in the figure.

QLoRA goes further: it stores the frozen base weights in 4 bits (a format called NF4) and trains LoRA adapters in 16-bit on top, which lets a 65-billion-parameter model be fine-tuned on a single 48 GB GPU.

Why it matters in practice. LoRA adapters are megabytes, not gigabytes. You can keep one per customer or task and hot-swap them on one shared base model, and training fits on far smaller hardware.

Chapter 5

Which adaptation should you use?

Everyday picture If a new employee doesn't know your product catalogue, you hand them the catalogue (retrieval); you don't send them back to school. If they know the facts but write emails in the wrong tone, you first give clearer instructions, then coach them, and only for a whole new profession do you retrain from scratch.

Approach Changes weights? Use when
Prompting / few-shot No behaviour or format change
RAG No knowledge that changes or must be cited
LoRA / QLoRA small adapters style, domain vocabulary, consistent output format
Full fine-tune yes, all rarely: major domain shift with lots of data

Figure 11 · Diagram

Reading it: start at the top. Knowledge problems exit immediately to retrieval. Behaviour problems climb a ladder of cost and only go as far as they must: prompt first, a LoRA adapter if the prompt can't make the behaviour consistent (or the prompt is too long and costly to send every time), and a full fine-tune only for a genuine domain shift with lots of data.

In code: choose_adaptation walks this flowchart, cheapest option first, and returns the approach it lands on.

The key line: fine-tuning teaches behaviour; RAG supplies knowledge. Most production systems end up as retrieval plus a well-built prompt.

Chapter 6

Distillation: the apprentice learns how the master hesitates

Soften the teacher's judgement yourself first; the temperature will then read as what it does.

Everyday picture A master chef tastes a sauce and says "mostly thyme, a bit of rosemary, definitely not mint". An apprentice who only hears "thyme" learns less than one who hears the whole judgement. Distillation trains a small, cheap student model to match a big teacher's full probability spread, not just its top answer.

Tiny worked example The teacher's raw scores (logits) for three answers are 2, 1, 0. Softmax at temperature T = 1 gives 0.665, 0.245, 0.090. Dividing the logits by T = 2 first gives softmax(1, 0.5, 0) = 0.506, 0.307, 0.186: flatter, so the student clearly sees that answer 2 is a much better runner-up than answer 3. That runner-up information is sometimes called "dark knowledge".

Figure 12 · Diagram

Reading it: both models see the same input. Each one's scores are softened by the same temperature, and the KL divergence measures how far the student's spread is from the teacher's. Only the student learns; the teacher is fixed. Typically this runs over large volumes of the exact traffic your product sees, so the student becomes an expert at your task.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
the teacher's logit (raw score) for answer i any real
temperature: divide scores by T before softmax; T > 1 flattens > 0
teacher and student probabilities at temperature T each sums to 1
Kullback-Leibler divergence: the extra surprise from believing q when the truth is p. Zero only when they match ≥ 0
natural logarithm
rescales the gradient, which softening shrinks by 1/T²
mix between matching the teacher and matching the true label 0 … 1
CE ordinary cross-entropy on the true label ≥ 0

In words: the student is penalised by how far its softened spread is from the teacher's, scaled by T², optionally mixed with the usual penalty on the true answer.

On the worked example: with teacher (0.5, 0.5) and student (0.9, 0.1), KL = 0.5·ln(0.5/0.9) + 0.5·ln(0.5/0.1) = −0.294 + 0.805 = 0.511. When the student matches the teacher exactly, KL = 0. If those are the two spreads at T = 2 and α = 1 (learn from the teacher alone), the loss is 2² × 0.511 = 2.04.

Level 3: in Python
import math
def soft_targets(z, T):
    # e^(z_i / T)
    exps = [math.exp(z_i / T) for z_i in z]
    # divided by Σ_j e^(z_j / T)
    return [e / sum(exps) for e in exps]
[round(p_i, 3) for p_i in soft_targets([2, 1, 0], T=2)]  # → [0.506, 0.307, 0.186]
# teacher, student
p, q = [0.5, 0.5], [0.9, 0.1]
KL = sum(p_i * math.log(p_i / q_i) for p_i, q_i in zip(p, q))
round(KL, 3)  # → 0.511
alpha, T = 1.0, 2
# the (1 − α)·CE term is zero at α = 1
round(alpha * T ** 2 * KL, 2)  # → 2.04

Figure 13 · Chart

T = 1 T = 2 T = 5 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 teacher probability Higher temperature reveals the runner-up ranking answer 1 (logit 2) answer 2 (logit 1) answer 3 (logit 0)

Raising temperature from 1 to 5 flattens the teacher's 0.665, 0.245, 0.090 towards even, revealing the ranking of wrong answers

Reading it: the same three teacher logits (2, 1, 0) shown at three temperatures. At T = 1 (left group) the top answer dominates. At T = 2 and T = 5 the bars even out and the ranking of the wrong answers becomes visible to the student. Too high a temperature and everything flattens to a uniform guess, so T is tuned (2 to 5 is common).

In code: soft_targets divides logits by T and applies softmax, kl_divergence measures how far apart two spreads are, and distillation_loss combines them with the T² scale and the optional cross-entropy on the true label.

Why it matters in practice. Distillation is often the biggest cost lever in production: a small student trained on a big model's outputs for one narrow task can be many times cheaper and faster with little quality loss on that task.

Test yourself

10 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: Why does a base model ramble instead of answering?Think it through, then reveal

A: Pretraining only teaches it to continue text. Answering questions in a helpful format is learned later, in SFT and preference tuning.

Question 2Q: What is the one code difference between pretraining loss and SFT loss?Think it through, then reveal

A: A mask. SFT computes the same next-token cross-entropy but averages it only over the reply tokens, so the prompt is read but not trained on.

Question 3Q: In RLHF, what does the reward model learn, and from what?Think it through, then reveal

A: A score for responses such that sigmoid(score gap) predicts which of two responses a labeler preferred. It learns from pairwise comparisons, which are far easier for people to give than perfect answers.

Question 4Q: How does DPO avoid a reward model?Think it through, then reveal

A: It treats the log-probability ratio between the policy and a frozen reference as an implicit reward, and applies the same pairwise loss directly to the policy. One model, one supervised-style training loop.

Question 5Q: What does β control in DPO?Think it through, then reveal

A: How tightly the policy is held to the reference model. It is the weight on the drift penalty in the objective DPO optimises, reward − β × KL(policy ‖ reference), so larger β keeps the policy closer to the reference and smaller β lets the preferences pull it further away. In the loss, a larger β makes each unit of log-ratio count for more, so pairs are satisfied with a smaller departure.

Question 6Q: Why is B initialised to zero in LoRA?Think it through, then reveal

A: So B·A = 0 and the adapted model starts exactly equal to the pretrained model. Training then learns only the change.

Question 7Q: How many parameters does LoRA rank 8 train on a 4096 × 4096 layer?Think it through, then reveal

A: 2 × 4096 × 8 = 65,536, about 0.39% of the 16.8 million in the full matrix.

Question 8Q: Does LoRA slow down inference?Think it through, then reveal

A: Not if you merge: add B·A into W once and serve the result. Unmerged adapters cost two thin extra matmuls, which is what lets you hot-swap adapters on one base model.

Question 9Q: A client wants the model to know their product catalogue, which changes weekly. Fine-tune or RAG?Think it through, then reveal

A: RAG. The knowledge changes and answers should cite sources; fine-tuning would be stale within a week and can't cite. Fine-tune only for behaviour the prompt can't make consistent.

Question 10Q: What does temperature do in distillation?Think it through, then reveal

A: It softens both distributions so the student learns the teacher's relative preferences among wrong answers, not just its top pick; the T² factor keeps gradient sizes comparable.

Primary sources

The papers behind this lesson

Brown et al., Language Models are Few-Shot Learners (GPT-3, 2020): Showed that scaling next-token pretraining produces a model that can follow instructions and examples given only in the prompt.

Read the annotated companion →The paper ↗

Kaplan et al., Scaling Laws for Neural Language Models (2020): , with Hoffmann et al., Training Compute-Optimal Large Language Models (2022): Measured how pretraining loss falls predictably with model size, data and compute, and how to balance them.

Read the annotated companion →The paper ↗

Ouyang et al., Training language models to follow instructions with human feedback (InstructGPT, 2022): Established the SFT, reward model and RLHF recipe behind chat assistants.

Read the annotated companion →The paper ↗

Rafailov et al., Direct Preference Optimization: Your Language Model is Secretly a Reward Model (2023): Showed preference tuning can skip the reward model and RL loop entirely.

Read the annotated companion →The paper ↗

Hu et al., LoRA: Low-Rank Adaptation of Large Language Models (2021): , with Dettmers et al., QLoRA (2023): Fine-tuned huge models by training small low-rank adapters beside frozen (and, in QLoRA, 4-bit) weights.

Read the annotated companion →The paper ↗

Hinton, Vinyals & Dean, Distilling the Knowledge in a Neural Network (2015): Trained small students on a large teacher's temperature-softened outputs.

Read the annotated companion →The paper ↗

Bai et al., Constitutional AI: Harmlessness from AI Feedback (2022): Replaced much human preference labelling with a model judging responses against written principles.

The paper ↗

Researcher's shelf

Further reading

  • Hugging Face PEFT documentation: https://huggingface.co/docs/peft/index
  • Hugging Face TRL documentation (SFT, DPO, reward modelling): https://huggingface.co/docs/trl/index

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.