rumblr Work in progressWIP

● The AI Primer · Lesson 33 · Embeddings, the centerpiece

Operating embeddings

model changes, jargon, and finding what broke

This lesson covers Model migrations, domain mismatch, measuring retrieval on its own

Members · open during launch 20 min7 figures and diagrams
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Every embedding model has its own vector space; never mix documents and queries from different models. Same-size models fail silently.
  2. Migrate blue/green: new index, dual-write, backfill, compare on a golden set, cut over only if better, keep the old index for rollback.
  3. Re-embedding cost is n × tokens: estimate hours and dollars before you start.
  4. General models miss company jargon; build an eval set from resolved logs, and fine-tune or add keyword search when needed.
  5. Triage RAG failures into retrieval vs. generation before fixing anything.

Level 2

How it works, from scratch

Level 2 makes each of those truths concrete, starting with two mapmakers.

The everyday picture. Two mapmakers each draw a map of the same city with their own grid. On one map, square (3, 7) is the train station; on the other, (3, 7) is a park. Both maps are fine, but you can't read a position off one map and look it up on the other.

Every embedding model is its own mapmaker. Upgrade the model and every document has to be re-plotted on the new map, and until that's done, the old map and the new one must never be mixed. That's the first of three operational truths this lesson makes concrete:

  1. Changing models means re-embedding everything, and doing it without downtime or a silent quality drop takes a plan.
  2. General models don't speak your company's jargon, so measure on your own questions and adapt when needed.
  3. When a retrieval-augmented system answers wrongly, find out which half failed: retrieval (the right page was never found) or generation (it was found and misused).

Chapter 1

Every model has its own space

Tiny worked example the word "vpn" embedded by two versions of the toy model (primer.common.embedder) with the same 64 dimensions: their cosine is about 0.06, essentially unrelated, though it's the same word. Now take 12 real questions with known answers (the "golden set" in primer.common.corpus) over 20 documents. Search v1 documents with v1 queries and the right answer is in the top 3 for 11 of 12 questions (0.92). Search the same v1 documents with queries from the new v2 model and it drops to 0.25, about what random ranking would give (3 of 20 documents shown, so 15%), and nothing crashes.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Example
Q the golden set of questions 12 questions
|Q| how many questions 12
q ∈ Q each question in turn
rel(q) the documents that truly answer q {it-004}
top_k(q) the k documents search returned 3 documents
∩, ≠ ∅ "they share at least one document"
[ … ] 1 if true, 0 if false
hit@k the share of questions with a right answer in the top k 0 to 1

In words: the share of golden questions for which at least one right document appears among the top k results. (With one relevant document per question, as here, this equals recall@k.)

On the example: 11 hits out of 12 = 0.92 with matching models; 3 out of 12 = 0.25 with mixed models.

In Python:

def hit(rel_q, top_q):
    # [rel(q) ∩ top_k(q) ≠ ∅]
    return 1 if rel_q & top_q else 0
hit({"it-004"}, {"it-002", "it-004", "fin-001"})  # → 1
# matching models: 11 of the 12 questions
hits = [1] * 11 + [0]
# (1/|Q|) Σ over q in Q
round(sum(hits) / len(hits), 2)  # → 0.92
# mixed models: 3 of the 12
hits = [1] * 3 + [0] * 9
round(sum(hits) / len(hits), 2)  # → 0.25

Figure 1 · Chart

v1 docs, v1 queries v2 docs, v2 queries v1 docs, v2 queries 0.0 0.2 0.4 0.6 0.8 1.0 recall@3 on the golden set Mixing two models' vectors fails silently 0.92 0.92 0.25 random ranking (3 of 20)

Recall@3 is 0.92 when documents and queries share a model, and falls to 0.25, near the 0.15 of random ranking, when v2 queries search v1 documents

Reading it: each bar is recall@3 on the golden set. The first two bars use one model for both documents and queries, v1 then v2: both work. The third bar searches v1 documents with v2 queries: recall falls to 0.25 (3 of the 12 questions), barely above the dashed line at 0.15, which is what random ranking would score (3 of 20 documents shown). A model with a different number of dimensions would at least crash (you can't dot a 128-number vector with a 64-number one); a same-size model fails silently, which is worse.

In code: golden_recall embeds the documents with one model and the questions with another and returns hit@k. VectorIndex is a flat index tied to one model version: VectorIndex.search embeds a text query, and VectorIndex.search_vector takes a vector that is already made.

Why it matters "we upgraded the embedding model and search got weird" is a classic incident. The cause is almost always documents and queries embedded by different models, typically because only new documents were re-embedded.

Chapter 2

Migrating without a bad day: blue/green

Everyday picture building a new bridge next to the old one. Traffic keeps using the old bridge while the new one is built and inspected. When the new bridge passes inspection, traffic is switched over, and the old bridge stays standing for a while in case something turns up.

Figure 2 · Diagram

Reading it: follow the states from the top. Users are served by v1 the whole time until cutover. Dual-writing from the very first step means documents added mid-migration land in both indexes, so the new one never falls behind. The shadow comparison is the gate: the switch only happens if v2 is at least as good on your own golden set (BlueGreenMigration refuses to cut over without it). The old index is kept after cutover, so rollback is flipping one pointer, not a multi-hour rebuild.

The "live alias" is a pointer: the search service asks for "the live index", and cutover or rollback just repoints it. Many vector databases support aliases for exactly this.

In code: each arrow of the diagram is one method: BlueGreenMigration.start, BlueGreenMigration.add (the dual write), BlueGreenMigration.backfill, BlueGreenMigration.shadow_compare, BlueGreenMigration.cutover and BlueGreenMigration.rollback.

Tiny worked example: what will re-embedding cost? 50 million documents of about 500 tokens each, an embedding throughput of 1 million tokens per second across your workers, and a price of $0.02 per million tokens (all three are inputs you replace with your own numbers):

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Example
n number of documents (or chunks) 50,000,000
t average tokens per document 500
n · t total tokens to embed 25,000,000,000
r throughput, tokens per second 1,000,000
3600 seconds per hour
p price per million tokens $0.02

In words: total tokens divided by throughput gives the time; total tokens in millions times the price gives the cost.

On the example: 25 × 10⁹ / (10⁶ × 3600) = 6.94 hours; 25,000 × $0.02 = $500.

In Python:

# documents, tokens per document
n, t = 50_000_000, 500
# tokens per second, dollars per million tokens
r, p = 1_000_000, 0.02
# total tokens
n * t  # → 25000000000
# hours
round(n * t / (r * 3600), 2)  # → 6.94
# dollars
round(n * t / 10**6 * p, 2)  # → 500.0

Figure 3 · Chart

1 0 5 1 0 6 1 0 7 1 0 8 1 0 9 documents (500 tokens each, log scale) 1 0 − 3 1 0 − 2 1 0 − 1 1 0 0 1 0 1 1 0 2 1 0 3 hours (log scale) Time to re-embed 100,000 tokens/s 1,000,000 tokens/s 10,000,000 tokens/s 1 0 5 1 0 6 1 0 7 1 0 8 1 0 9 documents (500 tokens each, log scale) 1 0 0 1 0 1 1 0 2 1 0 3 1 0 4 1 0 5 dollars (log scale) Cost to re-embed $0.01 per million tokens $0.02 per million tokens $0.13 per million tokens

Re-embedding time and cost both grow in step with corpus size: ten times the documents, ten times the hours and the dollars

Reading it: the horizontal axis is corpus size (log scale); the left panel shows hours at three throughputs, the right panel shows dollars at three prices. Both grow in straight lines on these log axes: ten times the documents, ten times the time and money. The money is usually modest; the time, the rate limits and the double storage during the migration are what need planning.

In code: reembed_estimate computes the tokens, hours and dollars for your own n, t, r and p.

Why it matters re-embedding is routine: new models, fine-tunes, chunking changes and bug fixes all require it. Versioned indexes, dual writes, a golden-set gate and a rollback path turn it from a risky event into a boring one.

Chapter 3

Domain mismatch: fluent, but not in your jargon

Everyday picture a new hire who speaks perfect English but doesn't yet know that "AP" means accounts payable or that "T&E" means travel and expenses. General embedding models are trained mostly on web text; your acronyms and product names are new words to them.

Tiny worked example four real-sounding employee questions in company jargon: "ap aging report", "hotspot from a client site", "sso lockout", "t-and-e submission deadline". The general toy model puts the right document first for 1 of 4. A version that has learned the four jargon words (DomainTunedEmbedder, standing in for a model fine-tuned on company pairs) gets 4 of 4.

Figure 4 · Chart

ap aging report hotspot from a client site sso lockout t-and-e submission deadline 0 2 4 6 8 10 12 14 rank of the right document (1 is best) Company jargon: general vs. tuned model general model tuned on the jargon

On company jargon the general model buries three of the four right answers, while the tuned model ranks every one first

Reading it: each pair of bars is one jargon question; the height is where the right document ranked (1 is best, shorter is better). The general model buries three of the four answers; the tuned model puts every one first. The one the general model gets right, "sso lockout", is saved by a word it does know ("lockout").

In code: rank_of_first_relevant finds where the right document lands for one question under one model: the height of each bar.

Figure 5 · Diagram

Reading it: the evaluation set comes from real usage, not invention. Resolved sessions give you (question, answer document) pairs for free; build_eval_set keeps only resolved ones and merges repeats. Test candidate models on that set, and when none is good enough, fine-tune on domain pairs (primer.ml.embeddings.contrastive) and add keyword search, which matches jargon exactly (primer.ml.embeddings.retrieval).

Why it matters public leaderboards measure public data. The only number that predicts your system's quality is recall on your own questions.

Chapter 4

Measure retrieval separately from generation

Everyday picture an open-book exam. A wrong answer has one of two causes: the right page wasn't in the book you brought (retrieval failure), or it was, and you misread it (generation failure). Studying harder fixes the second, not the first.

Tiny worked example three answered questions, checked by hand.

Retrieved (top 3) Truly relevant Answer used Verdict
it-002, fin-006, fin-001 it-004 it-002 retrieval failure: it-004 was never found
hr-002, hr-001, hr-005 hr-001 hr-002 generation failure: found but not used
it-001, it-002 it-001 it-001 correct

Figure 6 · Diagram

Reading it: one question splits every failure in two. If the right document wasn't retrieved, no prompt change can help, so work on retrieval. If it was retrieved and ignored, the problem is downstream, in the prompt or the context. Only the retrieval half can be measured without a model, which is why recall@k on a golden set is the first number to track.

Figure 7 · Chart

1 3 5 documents retrieved (k) 0 2 4 6 8 10 12 golden questions Which half failed? (generator uses the top document) correct generation failure retrieval failure

Retrieving more documents turns some retrieval failures into generation failures, and the error-code question stays wrong at every k

Reading it: each bar is the 12 golden questions, answered by a toy generator that always uses the top document, with k documents retrieved. Green is correct, red is a retrieval failure, orange is a generation failure. Retrieving more (larger k) turns some retrieval failures into generation failures: the right page is now in the book, but the reader still opened the wrong one. Watch the error-code question ("what does ERR-4012 mean"): its answer ranks only fourth, so it's a retrieval failure until k = 5, and even then the top document is the wrong one. Dense vectors blur exact identifiers; keyword search would put that page first.

In code: triage gives the verdict for one answered question, following the flowchart above; triage_golden_set runs every golden question through retrieval and a toy generator that answers from the top document.

Why it matters teams burn weeks tuning prompts for failures that were retrieval all along. Triage first, then fix the half that's broken.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: You need to switch embedding models on a 50-million-document index. Plan the migration.Think it through, then reveal

Create a versioned index for the new model and dual-write new documents to both. Backfill by re-embedding existing documents in restartable batches (estimate: 50M × 500 tokens = 25B tokens; at 1M tokens/s that's about 7 hours). Compare recall on a golden set in shadow; cut over the live alias only if the new model is at least as good; keep the old index for fast rollback, then retire it.

Question 2Q: Why can't you compare vectors from two different embedding models?Think it through, then reveal

Each model defines its own coordinate system. Even with the same number of dimensions, the same text lands in unrelated places, so similarity between the two spaces is meaningless.

Question 3Q: A general embedding model performs poorly on a company's internal documents. What helps?Think it through, then reveal

Build a small eval set from real queries and the documents that resolved them, test several models on it, add hybrid (keyword + dense) search for exact terms, and fine-tune an embedding model on domain pairs with hard negatives if needed.

Question 4Q: Your RAG system gives confident wrong answers. What's your first diagnostic step?Think it through, then reveal

Split retrieval from generation: for failing questions, check whether a relevant document was in the retrieved top k. If not, it's retrieval (fix search); if yes, it's generation (fix prompt and context). Track recall@k on a labeled set continuously.

Primary sources

The papers behind this lesson

Thakur et al., BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models (2021)

Showed that retrieval models trained on one domain often lose badly on others, which is why you evaluate on your own data.

The paper ↗
Muennighoff et al., MTEB: Massive Text Embedding Benchmark (2022)

Compared embedding models across many tasks and found no single model wins everywhere.

The paper ↗

Researcher's shelf

Further reading

  • Es et al., RAGAS: Automated Evaluation of Retrieval Augmented Generation (2023): https://arxiv.org/abs/2309.15217
  • Martin Fowler, BlueGreenDeployment: https://martinfowler.com/bliki/BlueGreenDeployment.html
  • MTEB leaderboard: https://huggingface.co/spaces/mteb/leaderboard

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.