rumblr Work in progressWIP

● The AI Primer · Lesson 12 · Part 1: how the model works inside

Reinforcement learning

learning from a score instead of an answer

This lesson covers Policy gradients from scratch, PPO, GRPO, reward hacking

Members · open during launch 47 min12 figures and diagrams9 interactive
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. RL learns from a score, not an answer: sample an action from the policy, get a reward, make rewarded actions more likely.
  2. REINFORCE: step along reward × ∇ log π(action); for softmax, ∇ log π is one-hot minus the probabilities. Unbiased but noisy.
  3. Baselines subtract the typical reward, turning rewards into advantages. Same average gradient, far less variance.
  4. PPO reuses each batch for several steps, clips the probability ratio to 1 ± ε so no step goes too far, and uses a value network as baseline plus a KL leash to a reference model.
  5. GRPO drops the value network: sample a group of answers per prompt and normalise rewards within the group. With a verifier as reward, it is how reasoning models are trained.
  6. Reward hacking: the policy optimises the reward you wrote, not the goal you meant. Defend with verifiable rewards, better reward models, a KL leash and a held-out check of the real goal.

Level 2

How it works, from scratch

Think of teaching a dog to sit. You can't show it the right answer: you can only wait for it to try something and give it a treat when the something was good. Over many tries the dog does more of what earned treats and less of what didn't. Nobody ever told it what "sit" means; it worked it out from a score.

That is reinforcement learning (RL). An agent (the dog, or a language model) takes an action (sits, or writes an answer), the world hands back a reward (a treat, or a score from a grader), and the agent adjusts itself so that rewarded actions become more likely. The agent's current habits, written as a probability for every action, are called its policy.

Compare this with ordinary supervised learning (primer.ml.losses), where every example comes with the correct answer attached. In RL there is no answer key, only a score after the fact. That is exactly the situation a language model is in once pretraining is over: for "prove this theorem" or "write a helpful reply" there is no single correct text to copy, but a checker or a judge can say how good an attempt was. RL is how the model learns from those judgements (primer.ml.training_stages places it in the training pipeline).

Figure 1 · Diagram

Reading it: follow the loop clockwise. The policy is a set of probabilities, and the action is sampled from it, so the agent sometimes tries things it isn't sure about. The environment is whatever judges the action; the agent can't see inside it. The only thing that comes back is one number, the reward, and the only thing the agent can do with it is nudge its own probabilities. Everything in this lesson is a better answer to one question: how exactly should that nudge be computed?

The dog from the first paragraph, with you holding the treats.

Chapter 1

A tiny worked example: three slot machines

Set a policy and pull the arms yourself first; the formula will then read as what it is.

The simplest RL problem is a row of slot machines, called a multi-armed bandit (a slot machine is a "one-armed bandit"). Machine A pays out 20% of the time, B 50% and C 80%, but the agent isn't told that. Each pull pays 1 or 0. The agent must find C by pulling and seeing what happens.

A policy here is three probabilities, one per arm. A policy that picks each arm a third of the time earns, on average, (0.2 + 0.5 + 0.8) / 3 = 0.5 per pull. A policy that always picks C earns 0.8. Learning means moving from the first policy to the second using nothing but the 1s and 0s.

This one number, the average reward a policy expects, is what RL maximises:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
one action: which arm to pull A, B or C
the policy's adjustable numbers (here, one score per arm, called logits) (0, 0, 0)
the policy: the probability of picking action , given . Here softmax of the logits 1/3 each
the average reward action pays (unknown to the agent) 0.2, 0.5, 0.8
add up over every action three terms
the expected reward: what the policy earns per pull, on average 0.5

In words: "the expected reward is each action's probability times its average payout, added up over the actions."

With the numbers: J = ⅓·0.2 + ⅓·0.5 + ⅓·0.8 = 0.5 for the uniform policy, and 1·0.8 = 0.8 for the policy that always pulls C.

Level 3: in Python
policy = [1/3, 1/3, 1/3]
# average payout of arms A, B, C
R = [0.2, 0.5, 0.8]
# J = Σ_a π(a) R(a)
round(sum(p * r for p, r in zip(policy, R)), 3)  # → 0.5
round(sum(p * r for p, r in zip([0, 0, 1], R)), 3)  # → 0.8

The agent can't compute J, because it doesn't know R. It can only sample: pull an arm, see a 1 or a 0. Every method below turns those samples into an estimate of which way to move θ to make J bigger. Trying an arm you're unsure of is called exploration; sticking with the best arm so far is exploitation. A sampled policy does some of both automatically, as long as no arm's probability has collapsed to zero.

In code: Bandit hides the win chances and pays out one pull at a time; expected_reward is the formula above.

Chapter 2

Policy gradients: do more of what worked (REINFORCE)

The lesson's one-step table, with the pull, the reward and the rate in your hands.

Everyday picture A football coach reviews the tape after a match. For every play that led to a goal, they tell the team "a bit more of that"; for plays that went nowhere, nothing. They don't need to know why the play worked. Repeat over hundreds of matches and the team drifts towards the plays that score.

Tiny worked example Start with logits (0, 0, 0), so each arm has probability ⅓. The agent pulls C and wins: reward 1. The rule, explained next, says: add to each logit learning rate × reward × (1 if it's the chosen arm, else 0, minus that arm's probability).

Arm Chosen? 1[chosen] − π × reward 1 × rate 0.5 New logit New probability
A no 0 − ⅓ = −0.333 −0.167 −0.167 0.274
B no 0 − ⅓ = −0.333 −0.167 −0.167 0.274
C yes 1 − ⅓ = +0.667 +0.333 +0.333 0.452

One lucky pull moved C from 33% to 45%. Had the pull paid 0, nothing would have moved. Had the agent pulled A and won (A wins sometimes too), A would have gone up instead. The rule is noisy, one pull at a time, but on average the arm that wins most gets pushed up most.

Figure 2 · Diagram

Reading it: the top path is acting: logits become probabilities, one action is sampled, and the environment pays a reward. The lower path is learning: from the action alone, work out which direction in logit space makes that action more likely (the ∇ log π box), then scale that direction by how good the outcome was. A big reward is a big step towards repeating the action; zero reward is no step.

The log-probability trick, decoded

We want the gradient of J: for each logit, how much J rises if the logit rises a little (see primer.notation for gradients from scratch). The difficulty is that J is an average over actions we can only sample. The trick rewrites the gradient as an average too, so a sample estimates it:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
"gradient with respect to θ": one slope per logit, collected into a vector 3 slopes
expected value: the average of the bracket when is sampled from the policy average over pulls
"drawn from"
the natural logarithm, the undo button for
the direction in logit space that makes action more likely, fastest for C
the -th logit (θ is the list of logits)
"partial derivative": the slope along one logit, holding the others still
1 if is the chosen action, else 0 (0, 0, 1)
the probability of action ⅓

In words: "the direction that raises expected reward is, on average, the direction that makes the sampled action more likely, weighted by the reward it earned. For a softmax policy, that direction is 'one for the chosen action, minus every action's probability'."

Why is this true? Because the slope of a probability equals the probability times the slope of its log (, the chain rule applied to log). So $\nabla J = \sum_a \nabla\pi(a) R(a) = \sum_a \pi(a) \nabla\log\pi(a) R(a)\pi(a)$ is an average over samples from π. The update is then plain gradient ascent: , with learning rate (primer.ml.optimizers).

With the numbers: at the uniform policy the true gradient is (−0.1, 0, +0.1): push C up, A down, leave B (which pays exactly the average) alone. One sample, "pulled C, got 1", estimates it as 1 × (−⅓, −⅓, ⅔). The step with α = 0.5 gives logits (−0.167, −0.167, 0.333) and probabilities (0.274, 0.274, 0.452), as in the table.

In Python:

import math
z = [0.0, 0.0, 0.0]
pi = [math.exp(z_k) / sum(math.exp(v) for v in z) for z_k in z]
# the true gradient: Σ_a π(a) R(a) (1[k=a] − π(k)), for each logit k
R = [0.2, 0.5, 0.8]
true_grad = [sum(pi[a] * R[a] * ((k == a) - pi[k]) for a in range(3)) for k in range(3)]
[round(g, 3) for g in true_grad]  # → [-0.1, 0.0, 0.1]
# one sample: pulled C (a = 2), reward 1
a, reward, alpha = 2, 1.0, 0.5
grad_log_pi = [(k == a) - pi[k] for k in range(3)]
[round(g, 3) for g in grad_log_pi]  # → [-0.333, -0.333, 0.667]
z = [z_k + alpha * reward * g for z_k, g in zip(z, grad_log_pi)]
[round(z_k, 3) for z_k in z]  # → [-0.167, -0.167, 0.333]
[round(math.exp(z_k) / sum(math.exp(v) for v in z), 3) for z_k in z]  # → [0.274, 0.274, 0.452]

This algorithm is called REINFORCE (Williams, 1992). Run it for 500 pulls and the policy finds arm C:

Figure 3 · Drawn from the lesson's code

0 100 200 300 400 500 pull 0.0 0.2 0.4 0.6 0.8 1.0 probability of picking the arm The policy finds the best arm arm A (wins 20%) arm B (wins 50%) arm C (wins 80%) 100 200 300 400 500 pull 0.3 0.4 0.5 0.6 0.7 0.8 0.9 reward, average of last 50 Reward rises as it learns best possible (always C) random pulling

REINFORCE on the three-armed bandit: the probability of arm C climbs from a third to about 0.96 within 500 pulls, and the average reward rises from 0.5 towards 0.8

Reading it: on the left, each line is one arm's probability over 500 pulls. All three start at ⅓. C's line (the 80% arm) climbs towards 1 while A and B sink; the wiggles are single lucky or unlucky pulls. On the right is the reward, averaged over the last 50 pulls. It starts near 0.5 (random pulling) and rises towards the dashed line at 0.8, the most any policy can earn. Nobody told the agent which arm was best: the 1s and 0s were enough.

Why it matters in practice. A language model is exactly this kind of policy, with a vocabulary of tokens as its arms, and one sampled answer is a string of sampled tokens. REINFORCE applies unchanged: sum the log probabilities of every token in the answer, and scale the gradient by the answer's reward. Every method below (PPO, GRPO) is REINFORCE with repairs.

In code: grad_log_prob is one-hot minus the probabilities, reinforce_step is one update, worked_reinforce_step is the table above, and train_reinforce runs the whole loop against a Bandit.

Chapter 3

Variance and baselines: grade on a curve

Drag the baseline across the two-arm example first; the variance numbers below will then be yours.

Everyday picture A teacher whose class all scores between 90 and 100 learns nothing by being told "you got a 92". What matters is whether 92 is above or below the class average. Raw scores that are all large and positive make every attempt look good; only the difference from typical tells you which way to go.

Tiny worked example Two arms, a 50/50 policy, and every pull pays a lot: arm 1 always pays 10, arm 2 always pays 12. Arm 2 is better, so the logit of arm 2 should rise. Look at the REINFORCE estimate for that logit:

Pulled Reward 1[arm 2] − π(arm 2) Estimate (no baseline) Estimate (baseline 11)
arm 1 10 0 − 0.5 = −0.5 10 × −0.5 = −5 (10 − 11) × −0.5 = +0.5
arm 2 12 1 − 0.5 = +0.5 12 × +0.5 = +6 (12 − 11) × +0.5 = +0.5

Without a baseline the estimate is −5 or +6 depending on the coin flip. It averages to +0.5, the right answer, but any single sample points the wrong way half the time, and violently. Subtract the average reward, 11, first and every sample says +0.5. Same average, no noise at all.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the baseline: any number that doesn't depend on which action was taken; usually the average reward 11
the advantage: how much better action did than typical −1 for arm 1, +1 for arm 2
everything else as in the REINFORCE formula above

In words: "scale each step by how much better than typical the action did, not by its raw reward."

Why is it allowed? Because the baseline's contribution averages to zero: $\mathbb{E}[, b, \nabla \log \pi(a)] = b \sum_a \nabla \pi(a) = b, \nabla \sum_a \pi(a) = b, \nabla 1 = 0$. Probabilities always add to 1, so pushing all of them up is impossible; the baseline only removes noise, never signal.

With the numbers: without a baseline the estimate's variance (the average squared distance from its mean, see primer.notation) is (25 + 36)/2 − 0.5² = 30.25. With b = 11 it is 0.

In Python:

# arm 1 pays 10, arm 2 pays 12; the 1[arm 2] − π(arm 2) factor for each pull
pulls = [(10, -0.5), (12, +0.5)]
def mean_and_variance(b):
    estimates = [(reward - b) * direction for reward, direction in pulls]
    mean = sum(estimates) / 2
    return mean, sum((e - mean) ** 2 for e in estimates) / 2
mean_and_variance(b=0)  # → (0.5, 30.25)
mean_and_variance(b=11)  # → (0.5, 0.0)

Figure 4 · Diagram

Reading it: the baseline sits between the reward and the update. Its job is to turn "how good was this?" into "how much better than usual was this?". The sign of the advantage now decides the direction of the step, so a below-average action is actively pushed down, even though its raw reward was positive.

To see it matter, give every arm of our bandit 5 extra points: rewards are now 5 or 6 instead of 0 or 1, and nothing about which arm is best has changed.

Figure 5 · Drawn from the lesson's code

no baseline baseline 5.5 1 0 − 1 1 0 0 1 0 1 variance of one-pull estimate Rewards of 5 or 6: gradient noise 20.36 0.15 0.0 0.2 0.4 0.6 0.8 1.0 final probability of the best arm, C no baseline baseline 20 runs each, 400 pulls

With rewards offset by 5, the gradient estimate's variance falls from about 20 to 0.15 with a baseline, and 20 training runs all find arm C with it, while without it most runs lock onto the wrong arm

Reading it: on the left, the variance of a one-pull gradient estimate at the starting policy, on a log scale: about 20 without a baseline and about 0.15 with one, over a hundred times smaller. On the right, each dot is one of 20 training runs (400 pulls each), placed at its final probability of picking C. With a running-average baseline (blue) every run ends near 0.95. Without one (red), the dots scatter to both ends: in most runs, early pulls of a mediocre arm paid 5 and were pushed up hard, and the policy committed before it ever learned C was better.

Why it matters in practice. Every practical policy-gradient method uses a baseline. PPO learns one with a second network, the value network (or critic), which predicts the expected reward from each state. GRPO, below, gets one for free by comparing several answers to the same prompt.

In code: gradient_estimate_stats computes the exact mean and variance of the estimate, sampled_gradient_variance measures it from real pulls, and train_reinforce subtracts a running average when baseline=True.

Chapter 4

PPO: take several steps, but never too far

The table of four cases, as one slider.

Everyday picture A chef tests a new recipe on one evening's diners. It would be wasteful to use their comments for just one small tweak, so the chef makes several rounds of changes from the same comment cards. But the further the recipe drifts from what the diners actually ate, the less their comments apply, so the chef caps each change: never more than 20% more or less of any ingredient per round.

For a language model, sampling answers is the expensive part (every answer is a full generation), so PPO (Proximal Policy Optimization) reuses each batch of answers for several gradient steps. It needs a way to tell how far the policy has moved since the batch was sampled, and a brake.

Tiny worked example The probability ratio compares the policy now with the policy that generated the sample. A ratio of 1.5 means the current policy is 50% more likely to produce that answer than when it was sampled. With a clip range ε = 0.2 the ratio is allowed to count only between 0.8 and 1.2:

Ratio ρ Advantage A ρ·A clip(ρ, 0.8, 1.2)·A min of the two What happened
1.1 +3 3.3 3.3 3.3 inside the band: plain REINFORCE
1.5 +2 3.0 1.2 × 2 = 2.4 2.4 good action already boosted enough: gain capped
0.5 −1 −0.5 0.8 × −1 = −0.8 −0.8 bad action already cut enough: capped
1.5 −1 −1.5 1.2 × −1 = −1.2 −1.5 bad action made more likely: full penalty
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
one sample in the batch; for a language model, one token of one answer row of the table
the state: what the policy saw before acting (the prompt plus the tokens so far) a prompt
the action taken (the token generated) an answer
the policy as it was when the batch was sampled, frozen
the policy now, after some steps on this batch
the probability ratio, new over old 1.5
the advantage of that sample +2
the clip range, typically 0.1 to 0.3 0.2
, but pushed back to or if it falls outside clip(1.5, 0.8, 1.2) = 1.2
the smaller of the two min(3.0, 2.4) = 2.4
the average over all samples in the batch
the objective PPO climbs

In words: "for each sample, take the ratio-weighted advantage, but once the ratio has moved more than ε from 1 in the direction the advantage wants, stop counting further movement; and always take the more pessimistic of the clipped and unclipped versions."

The slope of ρ·A is exactly REINFORCE's gradient scaled by ρ (because ), so inside the band PPO is REINFORCE with importance weighting. Outside it, the clipped term is flat: that sample stops pushing.

With the numbers: the four rows of the table, computed:

In Python:

def clip(x, lo, hi):
    return max(lo, min(x, hi))
def L_clip(rho, A, eps=0.2):
    return min(rho * A, clip(rho, 1 - eps, 1 + eps) * A)
[round(L_clip(rho, A), 2) for rho, A in [(1.1, 3), (1.5, 2), (0.5, -1), (1.5, -1)]]  # → [3.3, 2.4, -0.8, -1.5]
# past 1 + ε with a positive advantage, a higher ratio earns nothing more
L_clip(1.3, 2) == L_clip(1.6, 2)  # → True

Figure 6 · Diagram

Reading it: the outer loop is sampling, the expensive part: the frozen old policy writes a batch of answers and a value network turns their rewards into advantages. The inner loop reuses that batch for several passes. Each pass recomputes how far the policy has moved (the ratio), stops counting movement beyond the band (the clip), and applies the KL leash described below. When the passes are done, the updated policy becomes the new sampler and the cycle repeats.

To see the clip work, take one batch of 16 pulls from the bandit (C won all four of its pulls, A lost all four) and make 50 passes over it:

Figure 7 · Drawn from the lesson's code

0.4 0.6 0.8 1.0 1.2 1.4 1.6 1.8 probability ratio π_new / π_old −1.5 −1.0 −0.5 0.0 0.5 1.0 1.5 objective for one sample The clip: flat outside the band advantage +1 advantage -1 band 1 ± 0.2 0 10 20 30 40 50 pass over the same 16 pulls 1.00 1.25 1.50 1.75 2.00 2.25 2.50 largest ratio in the batch Reusing one batch 50 times no clip clip ε = 0.2

Two panels: the clipped objective is flat outside the 0.8 to 1.2 band on the side the advantage favours; over 50 passes the unclipped ratio climbs past 2.5 while the clipped one levels off near 1.24

Reading it: on the left is the objective for one sample as its ratio changes. For a positive advantage (blue), the line rises with the ratio until 1.2, then goes flat: no reward for pushing further. For a negative advantage (red), it goes flat below 0.8. The dashed lines are what REINFORCE would keep climbing. On the right, the largest ratio in the batch after each of 50 passes. Unclipped (red), the policy chases the same 16 pulls further every pass until C's probability is 2.5 times what it was, about 0.84 from sixteen pulls, which is wildly overconfident. Clipped (blue), it levels off near 1.24: the brake is not a hard wall (other samples can still nudge the policy), but it removes the incentive to overfit one batch.

In code: ppo_clipped_objective is the formula, clipped_policy_update makes several passes over one batch with or without the clip, and ppo_drift_experiment is the 50-pass comparison.

The KL leash: stay close to where you started

Everyday picture A dog on a long leash can explore, but it can't run off a cliff. In RL for language models, the leash ties the policy to a frozen copy of the model it started from, the reference model (usually the model after supervised fine-tuning).

Tiny worked example An answer earns reward 1.0. The policy now gives it probability 0.6; the reference gave it 0.3. With leash strength β = 0.1, the reward actually used for training is 1.0 − 0.1 × ln(0.6 / 0.3) = 1.0 − 0.1 × 0.693 = 0.931. The policy pays a small fee for having doubled that answer's probability.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the reward from the grader or reward model 1.0
the reward after the leash's fee 0.931
leash strength: how much a unit of drift costs 0.1
the frozen reference model's probability for the answer 0.3
how much more (positive) or less (negative) likely the policy makes this answer than the reference ln 2 = 0.693
KL divergence: the average of that log ratio over the policy's own answers; zero only when the two agree (decoded in primer.ml.training_stages)

In words: "each answer's reward is docked in proportion to how much more likely the policy has made it than the reference did; on average, that fee is β times the KL divergence between the two."

With the numbers: 1.0 − 0.1 × ln 2 = 0.931. Had the policy halved the answer's probability instead (0.15), the log ratio would be −0.693 and the reward would rise to 1.069: the leash pulls both ways.

Level 3: in Python
import math
R, beta = 1.0, 0.1
pi, pi_ref = 0.6, 0.3
# R' = R − β log(π / π_ref)
round(R - beta * math.log(pi / pi_ref), 3)  # → 0.931
round(R - beta * math.log(0.15 / pi_ref), 3)  # → 1.069

Why it matters in practice. This is the "penalty for drifting" in the RLHF loop of primer.ml.training_stages, and the same β appears in DPO. It stops the policy from forgetting fluent language while it chases reward, and it is the first line of defence against reward hacking, below.

In code: kl_penalised_reward is R′; clipped_policy_update adds the leash's gradient when given a reference policy and a β, and primer.ml.training_stages.kl_divergence computes the KL itself.

The leash's fee, with your own β, π and π_ref in it.

Chapter 5

GRPO: compare answers to the same question

Sample a group from the lesson's weak model yourself first; the advantage formula will then read as what it is.

Everyday picture Instead of hiring an examiner to predict how hard each exam question is, a teacher gives the same question to eight students and marks each answer relative to the others on that question. On an easy question, getting it right is expected and earns little credit; on a hard one, the only right answer stands out.

PPO's baseline comes from the value network, a second model, often as large as the policy, that has to be trained alongside it. GRPO (Group Relative Policy Optimization) throws the value network away. For each prompt it samples a group of answers and uses the group's own average as the baseline.

Tiny worked example The prompt is "3 + 4 =". The model samples four answers and a checker scores them 1 if the answer is 7, else 0.

Group rewards Mean Std Advantages
1, 0, 0, 1 0.5 0.5 +1, −1, −1, +1
1, 0, 0, 0 0.25 0.433 +1.73, −0.58, −0.58, −0.58
1, 1, 1, 1 1 0 0, 0, 0, 0

A lone right answer in a mostly wrong group earns a big advantage: it's rare, so it's strong evidence. A group that is all right (or all wrong) earns nothing: there is no contrast, so there is nothing to learn from.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the group size: answers sampled per prompt 4
which answer in the group 1 … 4
the reward for answer 1, 0, 0, 0
the group's average reward: the baseline 0.25
the group's standard deviation, the typical distance from the mean (square root of the variance) 0.433
answer 's advantage, shared by every token of that answer +1.73

In words: "an answer's advantage is how far its reward sits above the group's average, measured in units of the group's spread."

With the numbers: for (1, 0, 0, 0): mean 0.25, variance (0.75² + 3 × 0.25²) / 4 = 0.1875, std √0.1875 = 0.433, so the right answer gets 0.75 / 0.433 = 1.73 and each wrong one −0.25 / 0.433 = −0.58.

Level 3: in Python
import statistics
def advantages(r):
    mu, sd = statistics.mean(r), statistics.pstdev(r)
    return [round((r_i - mu) / sd, 2) if sd else 0.0 for r_i in r]
advantages([1, 0, 0, 1])  # → [1.0, -1.0, -1.0, 1.0]
advantages([1, 0, 0, 0])  # → [1.73, -0.58, -0.58, -0.58]
advantages([1, 1, 1, 1])  # → [0.0, 0.0, 0.0, 0.0]

The rest of GRPO is PPO: the same ratio, the same clip, the same KL leash to a reference model, averaged over the group. (Some implementations divide by the sample standard deviation, with G − 1, rather than the population one; the idea is identical.)

Figure 8 · Diagram

Reading it: both pipelines end in an advantage per answer, which feeds the same clipped update. PPO gets its baseline from a value network that must be trained, stored and run, which is roughly a second copy of the model. GRPO instead spends that compute on more answers per prompt and lets them grade each other. The verifier box can be any scorer, but GRPO shines when it is a program that checks the answer.

Verifiable rewards, and why GRPO trains reasoning

A verifiable reward comes from a check that can't be argued with: does the arithmetic equal 7, do the unit tests pass, does the proof check. No learned reward model, so no learned blind spots (see reward hacking, next). Here is GRPO on a toy "language model" that answers eight addition prompts with a single digit token. It starts out about 23% accurate and leans towards off-by-one mistakes.

Figure 9 · Drawn from the lesson's code

0 10 20 30 40 50 60 GRPO step (8 answers per prompt) 0.0 0.2 0.4 0.6 0.8 1.0 share GRPO with a verifier: 8 addition prompts chance of a correct answer prompts whose group all agreed (no signal)

GRPO on eight addition prompts: accuracy climbs from 23% to 99% in 60 steps, while the share of prompts whose whole group agreed, and so taught nothing, rises from a few percent to over 80%

Reading it: the blue line is the model's average chance of answering correctly, which climbs from 0.23 to 0.99 in 60 steps of 8 answers per prompt. The grey bars are the share of prompts whose group of 8 answers were all right or all wrong: a few percent at the start, over 80% by the end. They grow as the model masters the prompts: once every answer is right, the advantages are all zero and that prompt has nothing left to teach. Real GRPO training fights exactly this by filtering for prompts at the edge of the model's ability.

Why it matters in practice. This recipe, a verifier on the final answer plus GRPO, is how DeepSeek-R1-Zero learned to reason: rewarded only for correct final answers (and a required format), the model learned by itself to write longer chains of thought, to check its work and to back up from mistakes, because those behaviours raised the chance of a correct final answer. Each token in a long chain of thought shares its answer's advantage, so the whole chain is reinforced or discouraged together. See primer.ml.reasoning for what that training produces.

In code: group_advantages is the formula, verify is the checker, pretrained_logits is the weak starting model, and train_grpo runs the loop.

Chapter 6

Reward hacking: the score is not the goal

Scrub the training run before reading why the reward went up while the answers got worse.

Everyday picture A school pays tutors by the number of pages of homework feedback they write. Feedback gets longer, not better. An RL agent is the most literal-minded employee imaginable: it optimises the number you wrote down, not the thing you meant. When those two differ, it finds the difference. This is reward hacking (also called specification gaming), and it is Goodhart's law in code: when a measure becomes a target, it ceases to be a good measure.

Tiny worked example The goal is a good answer, and good answers here are about 4 sentences long. People rated answers of 1 to 4 sentences, and on those, longer really was better. A reward model fitted to their ratings learns a straight line: "each sentence is worth 0.19 points". Nobody ever rated a 10-sentence answer, so nobody told the reward model it was bad:

Sentences 1 2 3 4 6 8 10
True quality 0.44 0.75 0.94 1.00 0.75 0.00 −1.25
Reward model 0.50 0.69 0.88 1.06 1.44 1.81 2.19

The reward model's favourite answer, 10 sentences, is the worst one.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the answer's length in sentences 1 … 10
the reward model's score for a length- answer (the hat marks an estimate) 2.19 at n = 10
the rated examples: a length and its true quality (1, 0.44), …, (4, 1.0)
their averages (the bar means "mean") 2.5, 0.781
the fitted slope: points per extra sentence 0.1875
the fitted intercept 0.3125

In words: "the reward model is the straight line that best fits the ratings it saw: its slope is how length and quality moved together in the data, and it passes through the average point."

With the numbers: the deviations of length are (−1.5, −0.5, 0.5, 1.5) and of quality (−0.344, −0.031, 0.156, 0.219). Their products add to 0.9375 and the squared length deviations to 5, so w₁ = 0.1875 and w₀ = 0.781 − 0.1875 × 2.5 = 0.3125. At n = 10 it scores 2.19.

Level 3: in Python
n = [1, 2, 3, 4]
q = [1 - ((n_i - 4) / 4) ** 2 for n_i in n]
q  # → [0.4375, 0.75, 0.9375, 1.0]
n_bar, q_bar = sum(n) / 4, sum(q) / 4
w1 = sum((a - n_bar) * (b - q_bar) for a, b in zip(n, q)) / sum((a - n_bar) ** 2 for a in n)
w0 = q_bar - w1 * n_bar
(w0, w1)  # → (0.3125, 0.1875)
# the reward model's score for a 10-sentence answer, and its true quality
(w0 + w1 * 10, 1 - ((10 - 4) / 4) ** 2)  # → (2.1875, -1.25)

Figure 10 · Drawn from the lesson's code

2 4 6 8 10 answer length (sentences) −1.0 −0.5 0.0 0.5 1.0 1.5 2.0 score The reward model extrapolates; the truth turns over lengths people rated true quality reward model (a fitted line)

True quality peaks at 4 sentences and falls below zero past 8, while the reward model's straight line, fitted on lengths 1 to 4, keeps rising to 2.19 at 10 sentences

Reading it: the shaded strip is the only region anyone rated. Inside it, the reward model (red) and the truth (blue) agree on the direction: longer is better. Outside it, the reward model is extrapolating a straight line into territory it has never seen, while the truth turns over and dives. Every learned reward model has regions like this, and an optimiser is a machine for finding them.

Figure 11 · Diagram

Reading it: the solid path is how a reward is usually made: the goal is sampled by people, the ratings train a model, and that model, not the goal, is what the optimiser sees. The optimiser pushes the policy wherever the reward is highest, which, once the easy gains are taken, is wherever the proxy is most wrong. The dotted arrows are the two main defences: a leash that limits how far the policy can drift from where the ratings were collected, and a reward that has no learned gap to exploit.

Now train the starting model (which writes 2 or 3 sentences, a little too short) against each reward:

Figure 12 · Drawn from the lesson's code

0 50 100 150 200 250 300 training step 0.6 0.7 0.8 0.9 1.0 true quality of the policy Rise, then fall: reward hacking flawed reward, no leash flawed reward, KL leash β = 0.3 verifiable reward (the true goal) 1 0 − 4 1 0 − 3 1 0 − 2 1 0 − 1 1 0 0 1 0 1 KL from the reference (looser leash →) −1.0 −0.5 0.0 0.5 1.0 1.5 2.0 average score Best policy for each leash strength reward model's score true quality

Over 300 steps, optimising the flawed reward first raises true quality to 0.93 then drives it down to 0.75, below where it started, while a KL leash holds it near 0.84 and the verifiable reward climbs to 1.0; the right panel shows the best leashed policy's true quality peaking at a moderate KL and collapsing as the leash loosens

Reading it: on the left, true quality over 300 training steps. Against the flawed reward with no leash (red), quality rises at first, from 0.80 to 0.93, because lengthening a too-short answer genuinely helps. It rests on 5-sentence answers for a while, then the reward model's pull wins again: around step 150 the policy jumps to 6 sentences and quality falls to 0.75, below where it started, while the reward model's score keeps climbing. That rise-then-fall is the signature of reward hacking, and it is why "the reward went up" proves nothing. With a KL leash (β = 0.3, orange) the policy also overshoots a little, but the leash holds it at 0.84. Against the true, verifiable quality (blue) it reaches 1.0. On the right, the best policy for each leash strength, placed by how far it strays from the reference (its KL). Near zero KL it barely moves; at moderate KL true quality peaks; as the leash loosens further, the proxy score keeps rising and the true quality collapses below zero.

The right-hand panel uses a closed form: the best policy under a KL leash has an exact formula.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the policy that maximises (0.731, 0.269)
the reference model's probability for action (0.5, 0.5)
the reward for action (1, 0)
leash strength 1
a boost that grows with reward; a small β makes it enormous
a counter over every action, so adds up all of them
the total, so the probabilities add to 1 1.859

In words: "the best leashed policy starts from the reference and multiplies each action's probability by e to the power reward over β, then rescales so everything adds to 1."

With the numbers: two actions, reference (0.5, 0.5), rewards (1, 0), β = 1: weights 0.5 × 2.718 = 1.359 and 0.5 × 1 = 0.5, total 1.859, so π* = (0.731, 0.269). At β = 0.1 the boost is e¹⁰ ≈ 22,026 and π* puts 99.995% on the rewarded action; as β grows, π* returns to the reference.

Level 3: in Python
import math
ref, R = [0.5, 0.5], [1.0, 0.0]
def best_policy(beta):
    w = [p * math.exp(r / beta) for p, r in zip(ref, R)]
    return [round(w_a / sum(w), 5) for w_a in w]
best_policy(1.0)  # → [0.73106, 0.26894]
best_policy(0.1)  # → [0.99995, 5e-05]
best_policy(100.0)  # → [0.5025, 0.4975]

This is the same formula DPO starts from (primer.ml.training_stages): it is why β means the same thing in RLHF and DPO.

Why it matters in practice. Reward hacking shows up wherever RL does. A boat-racing game agent that learned to circle forever collecting bonus targets instead of finishing the race. RLHF'd chat models that learned long, flattering answers score well with raters (length bias and sycophancy). Coding models rewarded for passing tests that learned to edit or special-case the tests. The defences, strongest first:

  1. Verifiable rewards where the task allows: run the tests, check the answer. A check has no learned blind spot (though a buggy check does).
  2. Better reward models: rate the policy's current outputs and retrain, so the reward model sees the regions the policy is exploring; use ensembles; penalise known exploits such as length directly.
  3. A KL leash to keep the policy near the data the reward was fit on.
  4. Watch a held-out measure of the real goal (human review, a separate evaluation set, see primer.agents.evals) and stop when it turns down, even if the reward is still climbing.

In code: true_quality is the goal, fit_reward_model and proxy_reward are the flawed reward, optimise_lengths trains against either with an optional leash, and kl_regularised_optimum with leash_sweep gives the closed-form best policy for each β.

The closed form from the last formula, on one slider.

Test yourself

9 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1How does reinforcement learning differ from supervised learning?Think it through, then reveal

Supervised learning is given the correct output for every input and learns to copy it. Reinforcement learning is given only a score for the output it produced, so it must try things, see how they score and shift probability towards what scored well. It fits tasks where judging an answer is easy but writing the perfect one is not.

Question 2What is the log-probability trick, and why is it needed?Think it through, then reveal

The gradient of expected reward, Σ ∇π(a) R(a), can't be computed without knowing every action's reward. Rewriting ∇π = π ∇log π turns it into an average over actions sampled from the policy, E[R ∇log π(a)], so each sampled action and its reward give an unbiased estimate of the gradient.

Question 3Why does subtracting a baseline not change the expected gradient?Think it through, then reveal

Because E[b ∇log π(a)] = b ∇ Σ π(a) = b ∇ 1 = 0: probabilities always add to 1, so the baseline's push averages to nothing. It only removes the noise that comes from rewards being large or all the same sign.

Question 4In PPO, what is the probability ratio, and what does clipping it do?Think it through, then reveal

The ratio is the current policy's probability of a sampled action divided by the probability under the policy that sampled it; it measures how far the policy has moved on that sample. Clipping stops counting movement beyond 1 ± ε in the direction the advantage favours, so reusing a batch for several steps can't push the policy far from where the data came from.

Question 5Why does PPO's objective take the minimum of the clipped and unclipped terms?Think it through, then reveal

To stay pessimistic. Gains are capped once the ratio leaves the band, but if a step made a bad action more likely, the full penalty still applies, so the objective never rewards a harmful move.

Question 6How does GRPO get a baseline without a value network?Think it through, then reveal

It samples several answers to the same prompt and uses their mean reward as the baseline, dividing by their standard deviation to set the scale. Each answer is judged against its siblings, which saves training and serving a second model the size of the policy.

Question 7What happens in GRPO when every answer in a group gets the same reward?Think it through, then reveal

Every advantage is zero, so that prompt contributes no gradient. Prompts that are always solved or never solved teach nothing; learning comes from prompts at the edge of the model's ability.

Question 8What is reward hacking, and why does a KL penalty help against it?Think it through, then reveal

Reward hacking is the policy maximising the reward as written while the real goal gets worse, usually by finding inputs where a learned reward model is wrong. The KL penalty charges the policy for drifting from the reference model, which keeps it near the kind of outputs the reward model was trained on, where the reward is still trustworthy.

Question 9Why are verifiable rewards attractive for training reasoning?Think it through, then reveal

A program that checks the final answer (or runs the tests) has no learned blind spots to exploit and costs nothing to label, so RL can run for a long time against it without the reward drifting away from correctness.

Primary sources

The papers behind this lesson

Williams, Simple statistical gradient-following algorithms for connectionist reinforcement learning (Machine Learning, 1992)

Introduced REINFORCE, the log-probability policy gradient with a baseline.

Read the annotated companion →The paper ↗
Schulman et al., Proximal Policy Optimization Algorithms (2017)

Introduced the clipped probability-ratio objective that lets each batch be reused for several safe steps.

Read the annotated companion →The paper ↗
Ouyang et al., Training language models to follow instructions with human feedback (InstructGPT, 2022)

Used PPO with a per-token KL penalty to a reference model to tune a language model against a learned reward model.

Read the annotated companion →The paper ↗
Shao et al., DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (2024)

Introduced GRPO, replacing PPO's value network with group-relative advantages.

Read the annotated companion →The paper ↗
DeepSeek-AI, DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (2025)

Showed GRPO with rule-based, verifiable rewards alone can teach a model to produce long, self-checking chains of thought.

Read the annotated companion →The paper ↗
Gao, Schulman & Hilton, Scaling Laws for Reward Model Overoptimization (2022)

Measured how true quality rises and then falls as a policy is optimised further against a learned reward model.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Sutton & Barto, Reinforcement Learning: An Introduction (2nd edition, free online): http://incompleteideas.net/book/the-book-2nd.html
  • OpenAI, Spinning Up in Deep RL: https://spinningup.openai.com/
  • Andrej Karpathy, Deep Reinforcement Learning: Pong from Pixels: http://karpathy.github.io/2016/05/31/rl/
  • Lilian Weng, Policy Gradient Algorithms: https://lilianweng.github.io/posts/2018-04-08-policy-gradient/
  • Hugging Face TRL, GRPO Trainer: https://huggingface.co/docs/trl/grpo_trainer
  • Amodei et al., Concrete Problems in AI Safety (2016), section on reward hacking: https://arxiv.org/abs/1606.06565
  • Schulman et al., Proximal Policy Optimization Algorithms (2017): https://arxiv.org/abs/1707.06347
  • Shao et al., DeepSeekMath (2024), which introduces GRPO: https://arxiv.org/abs/2402.03300

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.