rumblr Work in progressWIP

● The AI Primer · Lesson 44 · Part 2: building systems people rely on

Retrieval-augmented generation (RAG), end to end

This lesson covers RAG end to end, with citations and access control

Free lesson 29 min11 figures and diagrams11 interactive
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. RAG retrieves relevant passages at question time and has the model answer from them with citations: knowledge that changes or must be cited, with no retraining.
  2. Quality is decided early: parsing (tables!), chunking with metadata, and hybrid retrieval with reranking.
  3. Filter by permissions before retrieval, never after generation.
  4. Upgrades: query rewriting, HyDE, multi-query, date filters, contextual retrieval, parent-child; agentic RAG for multi-part questions; GraphRAG for questions about connections.
  5. Debug by measuring retrieval (recall@k) first, then generation.

Level 2

How it works, from scratch

What follows builds every box of the pipeline from scratch, in the order a question travels through them, then the upgrades, then how to tell a retrieval failure from a generation failure.

An open-book exam with a librarian. Before you answer, the librarian fetches the few pages most likely to contain the answer; you answer from those pages and write down which page each fact came from. If the pages don't contain the answer, you say so instead of guessing.

Retrieval-augmented generation is that arrangement for a language model. The model's training knowledge is frozen and doesn't include your company's documents, so before each answer the system retrieves relevant passages from your documents and puts them in the prompt, and the model generates an answer from them, with citations. It's how you give a model knowledge that changes, or that must be cited, without retraining it.

Figure 1 · Diagram

Reading it: the top row runs once per document, whenever documents change; the bottom row runs for every question. Everything the answer can contain has to survive every box, left to right: a table mangled at Parse or a passage missing at Retrieve can't be fixed by a better prompt at Generate. That's why most RAG quality work happens in the first boxes, not the last.

In code: RAGIndex is the top row: it chunks every document and builds the keyword and dense indexes. answer is the bottom row, from filtered retrieval to verified citations, and returns a RAGResult.

The rest of this lesson walks the boxes in order, then covers the upgrades and how to debug a wrong answer.

The librarian's shelf, working: take a document off it and ask again.

Chapter 1

Parsing: where quality dies first

Repair the extraction yourself first; then read what each repair is for.

Everyday picture Photocopy a newspaper page and read it straight across: you get the first line of one column, then the first line of the next story, and nothing makes sense. Text extracted from PDFs has the same problems: running headers and page numbers mixed into the text, words broken across lines, and tables flattened into a row of numbers with no columns.

Worked example MESSY_PDF is two pages of a travel policy as a naive extractor returns it. clean_extracted_text repairs it in four steps:

Damage Before After
Header repeated on every page Northwind Ltd Travel Policy 2026 … CONFIDENTIAL removed (it appears on 2 of 2 pages)
Page-number footer Page 1 of 2 removed
Word hyphenated across a line break all em- / ployees travelling all employees travelling
Table flattened L1-L3 150 60 L4-L6 200 75 kept as rows, then one sentence per row

The table is the dangerous one. Flattened, "200" floats free of its grade. table_rows_as_sentences turns each row into a self-contained sentence, Grade L4-L6: hotel cap 200, meal cap 75., so any row can be retrieved on its own and still says what its numbers mean.

Figure 2 · Diagram

Reading it: each box removes one kind of extraction damage. The diamond is the important decision: prose and tables need different handling, because a table's meaning lives in its columns. Real pipelines add layout-aware parsers and OCR (optical character recognition: reading text from an image of a page) for scanned documents, with a quality check on the output.

Why it matters Enterprise documents are scanned, multi-column, full of tables that span pages. If parsing garbles the numbers, no retriever or model can recover them, and the failure looks like a model error.

In code: extract_table finds a table by its header row in the cleaned text and returns each row as a dict, ready for table_rows_as_sentences.

Chapter 2

Chunking, with metadata that travels

Everyday picture Index cards. Each card holds one idea, small enough to match a question precisely, and in its corner a label: which document, which date, who may read it.

Worked example With a 15-word limit, it-004 splits into two passages:

Passage id Text Updated Readers
it-004#0 ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date. 2026-04-10 everyone
it-004#1 Update AnyConnect to version 5.1 or later and reboot. 2026-04-10 everyone

(The limit is tiny because the toy documents are tiny; real systems use a few hundred tokens.) Split on the document's structure (sentences, paragraphs, sections), not at fixed character counts that cut a sentence in half. primer.ml.embeddings.retrieval compares chunking strategies in depth.

Why it matters Small chunks match precisely but can lose context (the second card doesn't say which error it fixes; see contextual retrieval below). The metadata is what makes permission and date filters possible later.

In code: chunk_document splits a document on sentence boundaries into Passages, and each Passage carries its id, title, department, date and readers.

Chapter 3

Retrieval: keyword and meaning, fused

Watch two searches disagree, and one list come out of them, before the formula.

Everyday picture Two librarians: one searches the catalogue for your exact words, the other understands what you mean even when you use different words. You take the books both of them rank highly.

  • Keyword search (BM25) scores passages by the question's words, giving more weight to rare words and less to long passages. It nails exact identifiers like ERR-4012 and misses synonyms.
  • Dense search compares embeddings (vectors where closeness means similar meaning; see primer.ml.embeddings). It finds "scam message in my inbox" → the phishing guide with no shared words, but blurs exact codes.
  • Hybrid search runs both and merges the two rankings with reciprocal rank fusion (RRF), which only needs the ranks, never the scores. BM25 scores and cosine similarities live on different scales, so adding them would be meaningless.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
one passage it-004#0
one of the rankings being fused (keyword, dense) 2 rankings
position of in ranking (1 = best); rankings that miss contribute nothing 1st, 3rd
a damping constant; 60 is the usual choice 60
add up over the rankings

In words: each ranking gives a passage a vote worth one over sixty-plus its position, and the votes are added.

On the worked example: a passage ranked 1st by dense search and 3rd by keyword search scores 1/61 + 1/63 = 0.0323. A passage ranked 1st by one list and absent from the other scores only 1/61 = 0.0164. Passages that both methods agree on rise to the top.

Level 3: in Python
k = 60
# 1st by dense, 3rd by keyword: Σ_i 1/(k + rank_i(d))
round(1 / (k + 1) + 1 / (k + 3), 4)  # → 0.0323
# 1st in one list, absent from the other
round(1 / (k + 1), 4)  # → 0.0164

Figure 3 · Chart

1 2 3 4 5 k (passages kept) 0.0 0.2 0.4 0.6 0.8 1.0 recall@k on the labelled questions Which retrieval finds the answer? keyword (BM25) dense hybrid (RRF) hybrid + rerank

Dense search plateaus at 92% recall because it never finds the error code, while hybrid reaches 100% by k = 2 and reranking lifts its top-1 from 75% to 92%

Reading it: the x-axis is how many passages you keep (k); the y-axis is recall@k, the share of the 12 labelled questions whose correct document is among those k. On this small set many questions reuse the documents' own words, so keyword search is already strong. Dense search plateaus at 92% because it never finds ERR-4012. Hybrid reaches 100% by k = 2 but only 75% at k = 1. Reranking the hybrid shortlist lifts the top-1 result to 92%. Run this measurement on your questions before choosing a method.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here
the labelled questions; is how many (12 here)
one question
1 if true, 0 if not

In words: the share of questions for which the right document made it into the top k.

On the worked example: dense search at k = 3 finds the right document for 11 of 12 questions (all but ERR-4012): 11/12 = 0.92.

Level 3: in Python
# 𝟙[...] for each of the 12 questions; ERR-4012 is the 0
found = [1] * 11 + [0]
# (1/|Q|) Σ over q in Q
round(sum(found) / len(found), 2)  # → 0.92

In code: RAGIndex.retrieve ranks the allowed passages with primer.ml.embeddings.retrieval.BM25, with dense similarity, or with both fused by primer.ml.embeddings.retrieval.reciprocal_rank_fusion, depending on the mode you ask for. recall_curve measures recall@k for each method.

Chapter 4

Reranking: read the question and each passage together

Everyday picture The librarians bring back twenty books quickly; an expert then reads your question next to each book and picks the best three.

A reranker (usually a cross-encoder: a model that reads the question and one passage together and outputs a relevance score) is far more precise than comparing precomputed vectors, and far too slow to run over every passage. So retrieval casts a wide, cheap net (here 8 candidates) and the reranker keeps the best few (here 3). This module reuses the toy cross-encoder from primer.ml.embeddings.retrieval.

In code: RAGIndex.rerank scores each (question, passage) pair with primer.ml.embeddings.retrieval.CrossEncoder and keeps the best few.

Chapter 5

Assemble, generate, cite, verify

Worked example "What does ERR-4012 mean?" becomes this context (each source escaped and tagged with its id; the question last; see primer.agents.context):

<sources>
<source id="it-004#1" title="Error ERR-4012: VPN tunnel failed" updated="2026-04-10">Update AnyConnect to version 5.1 or later and reboot.</source>
<source id="it-004#0" title="Error ERR-4012: VPN tunnel failed" updated="2026-04-10">ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date.</source>
<source id="it-002#2" title="Password policy" updated="2025-11-20">This policy applies to all employees and contractors.</source>
</sources>
<question>What does ERR-4012 mean?</question>

and the answer is "ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date. [it-004#0]". verify_citations then checks that every cited id was really in the context and that the cited passage supports the claim before it; a question the sources don't cover gets "I don't know based on the provided sources."

Figure 4 · Diagram

Reading it: time runs downward. Note the order: the permission filter runs at the index, before any passage is fetched; the model only ever sees three passages; and nothing reaches the user until the citations check out. Citations are what let a person verify an answer in seconds, which is most of what makes people trust the system.

In code: assemble_context builds the tagged sources and the question. generate sends them to the model with the answer-from-sources rules, and grounded_policy is the offline stand-in model that honours them. answer chains retrieve, rerank, assemble, generate and verify_citations.

The last step above, run on an answer you can edit.

Chapter 6

Permission-aware retrieval: filter before, never after

Try to read the forecast without being in finance, both ways.

Everyday picture A librarian checks your library card before fetching from the restricted archive. Checking it after you've read the document is pointless.

Each passage carries its access-control list (ACL: the groups allowed to read it), copied from the source system (SharePoint, Google Drive, a wiki). RAGIndex.retrieve drops every passage the user can't read before ranking, so restricted text can't reach the context, the model, or the answer.

Worked example "What is the Q3 revenue forecast?" (the forecast document is readable only by finance and exec):

Who asks How it's filtered Answer
finance user before retrieval "Q3 revenue is forecast at 41 million dollars… [fin-006#0]"
anyone else before retrieval "I don't know based on the provided sources."
anyone else after generation (citation removed) "Q3 revenue is forecast at 41 million dollars…"

Figure 5 · Diagram

Reading it: in the red box the model saw the restricted passage, so the number is already in its words, and deleting the citation marker afterwards leaves the fact behind. In the green box the passage never left the index. Permission changes in the source system must also sync to the index quickly, or a revoked user keeps access until the next re-index.

In code: post_generation_filter_answer is the wrong way, kept so the leak can be shown: it retrieves for every group, generates, then strips the restricted citation markers.

Chapter 7

Retrieval upgrades

Each fix switched off and on, on the question it was made for; the table below lists them all.

Each upgrade below is a working function, and each fixes a failure you can reproduce in demo():

Upgrade Everyday picture Before → after (top 3)
Query rewriting (rewrite_query) Repeating the earlier topic when you ask a follow-up "what if it fails?" after a VPN question: printer guide → VPN setup and ERR-4012
HyDE (hyde_search) Telling the librarian "I'm looking for a page that says something like…" "a weird message asking for my bank details": nothing relevant → phishing guide first
Multi-query (multi_query_search) Asking three librarians in three different phrasings "two step login setup": no MFA guide → MFA guide in top 3
Metadata filter (updated_after) Ignoring the out-of-date binder on the shelf superseded 2023 travel policy first → gone
Contextual retrieval (RAGIndex(contextual=True)) Writing the chapter title on every index card "how do I fix ERR-4012": fix sentence missing → found
Parent-child (retrieve_parents) Find the sentence, hand over the whole page a matching sentence → its full document

HyDE (Hypothetical Document Embeddings) deserves a picture, because it sounds backwards: you ask a model to invent an answer, then search with it.

Figure 6 · Diagram

Reading it: questions and answers are written differently. The user's words ("weird", "bank details") appear nowhere in the phishing guide, but the drafted answer uses the vocabulary answers use ("phishing", "report", "suspicious"), so its embedding lands next to the real passage. The draft may be wrong in its details; that's fine, because it's only used to search and is never shown to the user.

Contextual retrieval fixes the "second index card" problem from the chunking section: before indexing, each passage gets a short line saying where it comes from. Here that's just the title and department; Anthropic's version asks a model to write one sentence situating the chunk in its document.

Figure 7 · Chart

0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 rank of the passage containing the fix (lower is better; 21 = not in top 20) how do I fix ERR-4012 ERR-4012 what should I do resolve ERR-4012 on my laptop Contextual retrieval plain chunks chunks with a context line

Without a context line the fix passage misses the top 20 for all three phrasings; with the title prefixed it ranks 2nd, 2nd and 4th

Reading it: three ways of asking how to fix ERR-4012; each pair of bars is the rank of the passage "Update AnyConnect to version 5.1 or later and reboot." (shorter is better, and 21 means not in the top 20). Without context the passage never mentions ERR-4012, and for all three phrasings it doesn't make the top 20 at all. With the title prefixed it ranks 2nd, 2nd and 4th: in the top five every time, though "resolve ERR-4012 on my laptop" still ranks three passages from the laptop guide above it. A context line turns "never found" into "found", and ranking the right passage first is then the reranker's job.

Chapter 9

GraphRAG: questions about connections

Everyday picture A detective's corkboard: photos of people and places joined by string, each string labelled with the note that established it.

Some questions aren't answered by any single passage: "what depends on two-factor authentication?" or "what are the main themes across these contracts?". GraphRAG first extracts entities (named things: systems, people, products) and the relationships between them into a graph, then answers by walking it, or by summarizing groups of connected entities. build_graph does the smallest honest version: two entities are linked when one sentence mentions both, and each link remembers which document said so.

Figure 9 · Diagram

Reading it: each circle is an entity found in the knowledge base; each line is a sentence that mentioned both ends, labelled with its source document. "What depends on two-factor?" is answered by reading the neighbours of one circle, with citations for free. Real GraphRAG uses a model to extract typed relations ("requires", "replaces") and to write a summary for each cluster, which is what makes broad, whole-collection questions answerable.

In code: communities groups the graph from build_graph into connected clusters of entities, the "themes" a real GraphRAG system would summarize.

Chapter 10

Debugging a confident wrong answer

Everyday picture A student gets an exam question wrong. Either the librarian brought the wrong book, or the student misread the right one. The fixes are completely different, so find out which before changing anything.

Worked example debug_wrong_answers runs the 12 labelled questions and classifies each: retrieval miss (no relevant document in the top 3: no prompt can fix it), generation miss (the right document was there but the answer didn't use it), or ok. With dense-only retrieval, "what does ERR-4012 mean" is a retrieval miss and recall@3 is 0.92; switching to hybrid fixes it (recall@3 = 1.00). What's left are generation misses. One of them, "per-diem for meals when traveling", is answered from the superseded 2023 policy: a data problem that a date filter fixes. The other, "enroll in MFA", is a limit of this lesson's stand-in model rather than of RAG: it matches words literally, so "enroll" misses the passage's "enrolling", and with only "MFA" in common it declines to answer. A real model would read it-006 and answer.

Figure 10 · Diagram

Reading it: start at the top and answer each question with a measurement, not a guess. Most teams jump straight to the bottom-right box and edit the prompt; the diagram says to rule out the three earlier failure points first, because they're more common and a prompt can't fix them.

Figure 11 · Chart

0 2 4 6 8 10 12 labelled questions (top 3 passages) bm25 dense hybrid Where do wrong answers come from? ok generation miss retrieval miss

Dense search has the only retrieval miss; switching to hybrid removes it, and the wrong answers left are generation misses

Reading it: each bar is one retrieval method over the 12 labelled questions, split into ok (green), generation misses (orange) and retrieval misses (red). Dense search has a retrieval miss (the error code); hybrid removes it. The orange that remains is where to look next, and it's not retrieval.

The diagnosis above, on all twelve labelled questions: change the retrieval and watch the misses move.

Test yourself

5 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Your RAG system gives confident wrong answers. Debug it step by step.Think it through, then reveal

Take the failing questions (and build a labelled set if you don't have one). First, is the right passage in the index at all? If not, it's ingestion: parsing, chunking or permission sync. Second, is it in the top k? Measure recall@k; if not, fix retrieval: hybrid search, a reranker, query rewriting or HyDE, metadata filters, contextual chunks, domain-tuned embeddings. Third, did it survive into the final context? If not, fix reranking or the context budget. Only then look at generation: instructions to answer only from sources and to decline otherwise, citation checks, stale or contradictory sources. Add every case to the golden set.

Question 2Why does hybrid search beat pure vector search on enterprise data?Think it through, then reveal

Enterprise questions are full of exact identifiers (error codes, product names, ticket numbers) that embeddings blur, and full of paraphrases that keyword search misses. Hybrid gets both, and RRF merges the rankings without having to reconcile incompatible score scales.

Question 3How do you make RAG respect document permissions?Think it through, then reveal

Copy each document's access-control list onto every chunk at ingestion, filter candidates by the user's groups before ranking, and sync permission changes from the source system promptly. Never filter after generation: the model has already used the restricted text.

Question 4When would you use agentic RAG instead of a single retrieval step?Think it through, then reveal

For multi-part or vague questions, and when the model needs to judge whether results are good enough and search again. It costs more calls and latency, so keep single-shot retrieval for simple lookups.

Question 5RAG or fine-tuning for a company knowledge assistant?Think it through, then reveal

RAG for knowledge: it updates by re-indexing, cites sources, and respects permissions. Fine-tuning teaches behaviour and format, not facts that change. Most systems are RAG plus a good prompt (see primer.ml.training_stages).

Primary sources

The papers behind this lesson

Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020), Coined the term and showed that pairing a generator with a retriever over a document index beats a model relying on its weights alone for knowledge-heavy questions.

Read the annotated companion →The paper ↗

Gao, Ma, Lin & Callan, Precise Zero-Shot Dense Retrieval without Relevance Labels (2022), Introduced HyDE: search with an embedding of a model-written hypothetical answer.

Read the annotated companion →The paper ↗

Edge et al., From Local to Global: A Graph RAG Approach to Query-Focused Summarization (2024), Builds an entity graph and community summaries so that whole-collection questions can be answered.

The paper ↗

Researcher's shelf

Further reading

  • Anthropic, Introducing Contextual Retrieval: https://www.anthropic.com/news/contextual-retrieval
  • Anthropic citations docs: https://docs.claude.com/en/docs/build-with-claude/citations
  • Microsoft GraphRAG documentation: https://microsoft.github.io/graphrag/
  • RAGAS documentation (RAG evaluation): https://docs.ragas.io/
  • The retrieval lesson in this repo, with BM25, RRF, cross-encoders and chunking in depth: primer.ml.embeddings.retrieval

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.