At a glance
Key takeaways
- Contrastive training pulls matching pairs together and pushes non-matches apart; the loss (InfoNCE) is softmax cross-entropy over the batch.
- In-batch negatives are free: every other answer in the batch.
- Negatives only teach the distinctions they contain. Hard negatives (same topic, wrong answer) are what teach retrieval rather than topic matching.
- Temperature sharpens the softmax so small cosine gaps count; typical τ is 0.01 to 0.1.
- CLIP applies the same loss in both directions between an image encoder and a text encoder, which gives one shared space and zero-shot classification.
Level 2
How it works, from scratch
A teacher has a stack of question cards and a stack of answer cards. She lays out a few questions, deals out all the answers face up, and asks the student to pair each question with its answer. Every wrong pairing, she corrects. Over many rounds the student learns what makes an answer fit a question.
Then she gets sneaky. Alongside each right answer she slips in a look-alike wrong card: same subject, wrong answer. "How do I reset my password?" now faces both the reset steps and the password policy. A student who only ever saw easy wrong cards (answers about printers or holidays) would happily pick the policy card because it says "password". The look-alikes force the student to learn the difference that actually matters.
That's contrastive training, and it's how nearly every modern embedding
model is taught. The student is the model; "pairing" means making a
question's vector point the same way as its answer's vector (see
primer.ml.embeddings.similarity); the wrong cards are negatives; the
look-alikes are hard negatives.
Chapter 1
A tiny worked example: one question, two answer cards
A question's vector is q = (1, 0). The right answer is p₊ = (1, 0) and a wrong one is p₋ = (0, 1). All three have length 1, so a dot product is a cosine.
- Score each card with the dot product: q·p₊ = 1, q·p₋ = 0.
- Divide by the temperature τ (tau), a sharpness dial explained below. With τ = 1 the scores stay (1, 0).
- Softmax (from
primer.ml.attention): exponentiate and share out. e¹ = 2.718 and e⁰ = 1, so the right card gets 2.718 / 3.718 = 0.731. - Loss = −ln 0.731 = 0.313. It would be 0 if the right card got 100%.
At τ = 0.1 the scores become (10, 0), the right card gets 99.995%, and the loss is 0.0000454. Same vectors, far more confident: that's what the temperature does.
Figure 1 · Diagram
flowchart LR Q["Question: reset my password"] --> E1[Encoder] P["Right answer: reset steps"] --> E2[Encoder] N["Look-alike: password policy"] --> E3[Encoder] E1 --> PULL[Pull these two<br/>vectors together] E2 --> PULL E1 --> PUSH[Push these two<br/>vectors apart] E3 --> PUSH
Chapter 2
The math: InfoNCE, the loss behind it
In practice the teacher deals a whole batch. B questions sit in rows, all the answer cards in the batch sit in columns, and every card that isn't a question's own answer is a free negative for it: in-batch negatives.
Figure 2 · Diagram
flowchart LR
subgraph Batch["Similarity table for a batch of 3"]
direction TB
r1["q₁: ✔ ✗ ✗ | ✗"]
r2["q₂: ✗ ✔ ✗ | ✗"]
r3["q₃: ✗ ✗ ✔ | ✗"]
end
Batch --> SM["softmax along<br/>each row"] --> L["loss: −log of the<br/>✔ cell's share"]
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| B | number of questions in the batch | 8 in this lesson |
| M | number of answer cards in the batch (B, plus any hard negatives) | 8 or 16 here |
| i | which question (row) | 1 to B |
| j | which answer card (column) | 1 to M |
| qᵢ | the i-th question's vector, unit length | d numbers (d = 16) |
| pᵢ | question i's own answer vector (the ✔) | d numbers |
| pⱼ | the j-th answer card in the batch | d numbers |
| · | dot product; equals cosine because the vectors are unit length | −1 to 1 |
| τ (tau) | temperature: divides every score; small τ sharpens the softmax | 0.01 to 1; 0.1 here |
| exp | e raised to the power | > 0 |
| Σⱼ | add up over all M answer cards | |
| the fraction | softmax share of the right card | 0 to 1 |
| log, −(1/B) Σᵢ | take −log of each row's share, then average over the rows | L ≥ 0 |
In words: for every question, measure what share of its softmax goes to its own answer, take the negative log (big when the share is small), and average over the batch.
On the example: B = 1, M = 2, τ = 1: L = −log(e¹ / (e¹ + e⁰)) = −log 0.731 = 0.313.
In Python:
import math
def dot(a, b):
return sum(a_i * b_i for a_i, b_i in zip(a, b))
def info_nce(q, p, tau):
B = len(q)
total = 0
for i in range(B):
# exp(q_i · p_j / τ) for every card j
exps = [math.exp(dot(q[i], p_j) / tau) for p_j in p]
# log of the ✔ card's share
total += math.log(exps[i] / sum(exps))
# -(1/B) Σ_i
return -total / B
# B = 1 question
q = [(1, 0)]
# M = 2 cards; card 0 is the question's own answer
p = [(1, 0), (0, 1)]
round(info_nce(q, p, tau=1), 3) # → 0.313
print(f"{info_nce(q, p, tau=0.1):.7f}") # → 0.0000454
The name InfoNCE comes from "noise-contrastive estimation": telling the true
pair apart from noise. It's the softmax cross-entropy loss from
primer.ml.losses, where the "classes" are the answer cards in the batch.
The model trained here is a deliberately tiny bi-encoder (the same
encoder applied separately to questions and answers): count the words in a
text (a bag of words), multiply by one learned matrix W, and scale the
result to length 1. _contrastive_grads and _normalize_backward derive the
gradient (the direction in which each number in W should move to lower
the loss) by hand, and encoder_gradient_check confirms it against a
brute-force numerical estimate.
In code: info_nce_loss computes L for a batch, treating any rows of P
beyond B as extra hard negatives. BagOfWords turns text into word counts,
BiEncoder holds the learned matrix W, and BiEncoder.encode maps texts to
unit vectors.
Why it matters the embedding models behind search and RAG (Sentence-BERT, DPR, E5 and their successors) are trained with exactly this loss on millions of (question, answer) pairs. When a retrieval system confuses "reset my password" with "password policy", this is the training signal that was missing.
Chapter 3
Temperature: the sharpness dial
Everyday picture grading on a curve. With a gentle curve (high τ), a slightly better answer gets slightly more credit. With a steep curve (low τ), the best answer takes nearly all the credit, so small differences in score become big differences in outcome.
Tiny example the right card has cosine 0.8 and three wrong cards have 0.6 each.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Range |
|---|---|---|
| share₊ | the softmax share that goes to the right card | 0 to 1 |
| 0.8, 0.6 | cosine of the right card and of each wrong card | −1 to 1 |
| 3 | the number of (identical) wrong cards | |
| τ (tau) | temperature: every cosine is divided by it | 0.01 to 1 in practice |
| e^x | e ≈ 2.718 raised to the power x | > 0 |
In words: the right card's share is its exponentiated, temperature-scaled score divided by the total over all cards.
On the example: at τ = 1, 2.2255 / (2.2255 + 3 × 1.8221) = 0.289; at τ = 0.05 the scores become 16 and 12, and 1 / (1 + 3e⁻⁴) = 0.948.
In Python:
import math
def share_plus(tau):
# e^(0.8/τ)
right = math.exp(0.8 / tau)
# three wrong cards, e^(0.6/τ) each
wrong = 3 * math.exp(0.6 / tau)
return right / (right + wrong)
round(share_plus(1), 3), round(share_plus(0.05), 3) # → (0.289, 0.948)
Figure 3 · Chart
With the right card 0.2 ahead in cosine it gets under a third of the softmax share at temperature 1 and about 95% at temperature 0.05
In code: positive_probability computes share₊ for given cosines and τ;
this curve is that function swept across τ.
Why it matters too high a temperature and the model can't become confident; too low and a few hard pairs dominate training and it becomes unstable. It's one of the handful of settings that really matters when fine-tuning an embedding model.
Chapter 4
Hard negatives: the experiment
Everyday picture a driving test on an empty road teaches you to steer. It teaches nothing about merging, because merging never came up. Negatives only teach the distinctions they contain.
The setup (train_bi_encoder): questions come in two intents for each
topic: "how do I fix it?" (answered by a how-to card: "restart, reinstall
and follow the setup guide") and "what are the rules?" (answered by a policy
card: "approval, compliance and usage limits"). The question and answer
cards share almost no words, so matching intent must be learned ("fix"
goes with "restart"), not read off the overlap. Each batch holds one intent
across different topics, so its in-batch negatives differ from the right
card only by topic. We train two identical models; the only difference is
that one also gets each question's same-topic, other-intent card as a hard
negative. Then we test on four topics neither model has seen, including
"password".
Figure 4 · Diagram
flowchart TD D[Training pairs] --> A[Batches: one intent,<br/>8 different topics] A --> M1[Model 1: in-batch negatives only<br/>wrong cards differ by topic] A --> H[Add each question's<br/>look-alike card] H --> M2[Model 2: plus hard negatives<br/>wrong cards also differ by intent] M1 --> T[Test on unseen topics:<br/>right card vs. look-alike] M2 --> T
Figure 5 · Chart
Without hard negatives the model stays mostly below a coin flip on look-alike cards and ends near 60%, while with hard negatives it reaches 100% within ten steps
Figure 6 · Chart
With in-batch negatives only, held-out questions and cards group by topic; with hard negatives they split by intent, each question beside the card that answers it
primer.notation).
Circles are questions and squares are answer cards; blue means how-to and
orange means policy. With in-batch negatives only (left), the points group
by topic, and how-to and policy questions sit on top of each other. With
hard negatives (right), they split by intent, and each question sits next
to the kind of card that actually answers it.Figure 7 · Chart
On the unseen password topic, the in-batch-only model scores both cards almost alike, while the hard-negative model lights up the reset card for how-to questions and the policy card for policy questions
In code: make_pairs builds the (question, right card, look-alike)
triples; look_alike_accuracy runs the exam on unseen topics, and
training_curve records that score at checkpoints during training.
Why it matters hard negatives are the biggest single driver of retrieval quality. In practice you mine them: search your corpus with BM25 or the current model, and take high-ranking results that are not the right answer.
Chapter 5
Fine-tuning on your own domain
Figure 8 · Diagram
flowchart LR L[Logs: real questions +<br/>the document that resolved each] --> P[Positive pairs] P --> MINE[Mine hard negatives:<br/>top results that are wrong] MINE --> T[Fine-tune with InfoNCE,<br/>in-batch + hard negatives] T --> E["Evaluate recall@k<br/>on a held-out set"] E -->|better than base model| D[Re-embed corpus, deploy] E -->|not better| P
primer.ml.embeddings.operations).A few thousand good pairs is usually enough to adapt a general model to
legal, medical or internal jargon. Libraries such as sentence-transformers
implement this loss as MultipleNegativesRankingLoss.
Test yourself
5 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Q: How are embedding models trained?Think it through, then reveal
Contrastively, on (query, relevant passage) pairs: an encoder embeds both, and InfoNCE rewards the pair's similarity over the similarity to other passages in the batch (in-batch negatives) and to deliberately chosen hard negatives.
Question 2Q: Why do hard negatives matter so much?Think it through, then reveal
Random negatives are usually about something else, so a model can beat them by topic matching alone. Hard negatives share the topic but don't answer the question, forcing the model to learn the fine distinction that retrieval actually needs.
Question 3Q: What does the temperature do in InfoNCE?Think it through, then reveal
It divides the similarities before softmax. A small temperature makes the softmax sharp, so small cosine differences produce large differences in probability and gradient. Too small makes training unstable; too large makes the model unable to become confident.
Question 4Q: How does CLIP put images and text in the same space?Think it through, then reveal
It trains an image encoder and a text encoder together on image-caption pairs with a symmetric contrastive loss: each image must pick its caption from the batch and each caption must pick its image. Matching pairs end up close in one shared space.
Question 5Q: How would you adapt an embedding model to a company's jargon?Think it through, then reveal
Collect real (query, correct document) pairs from logs, mine hard negatives from current search results, fine-tune with in-batch plus hard negatives, evaluate recall@k on a held-out set against the base model, then re-embed the corpus with the new model.
Primary sources
The papers behind this lesson
Named and popularised the InfoNCE loss: pick the true sample out of a set of negatives.
The paper ↗Turned BERT into a bi-encoder that produces one comparable vector per sentence, making semantic search with transformers fast.
Read the annotated companion →The paper ↗Showed a question/passage bi-encoder trained with in-batch negatives plus BM25-mined hard negatives beats keyword search for open-domain QA.
Read the annotated companion →The paper ↗Trained image and text encoders with a symmetric contrastive loss on 400 million pairs, giving zero-shot image classification.
Read the annotated companion →The paper ↗Showed contrastive learning of sentence embeddings works even with dropout noise as the only "augmentation", and with NLI pairs as hard negatives.
The paper ↗Researcher's shelf
Further reading
- sentence-transformers, training overview: https://www.sbert.net/docs/sentence_transformer/training_overview.html
- Wang et al., Text Embeddings by Weakly-Supervised Contrastive Pre-training (E5, 2022): https://arxiv.org/abs/2212.03533
- Muennighoff et al., MTEB: Massive Text Embedding Benchmark (2022): https://arxiv.org/abs/2210.07316
- MTEB leaderboard: https://huggingface.co/spaces/mteb/leaderboard
- Lilian Weng, Contrastive Representation Learning: https://lilianweng.github.io/posts/2021-05-31-contrastive/
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.