RAG Architecture & Implementation
Why and when to use RAG, how its stages fit together, where wrong answers come from, and what RAG does and does not guarantee.
SPACED REPETITION Β· 15 practice questions
Make this lesson stick.
Try 3 questions now. No account needed. Sample answers aren't saved.
or sign in to practice all 15The same wrong answer, four weeks in a row
On Monday, February 2, 2026, an employee asks the company assistant:
How many unused vacation days can I carry over into next year?
Version 0 is a plain LLM call. It says that many companies allow about five days and suggests checking the handbook. The answer is fluent, plausible and useless. The leave policy is a private document the model never saw in training, so it produced the most likely-sounding answer, not your answer.
Version 1 adds retrieval. The HR documents are split into chunks and indexed. At question time the system finds the most relevant chunks, pastes them into the prompt, and asks the model to answer from them and cite them. That is retrieval-augmented generation (RAG). It replied:
You can carry over up to 5 days [Employee Handbook 2024, Β§4.2].
It has a citation, a confident tone and the wrong number. The Leave Policy 2026, effective January 1, raised the cap to 10 days and added a January 31 deadline for requests. Over the next weeks the team changed the assistant three times, and the answer stayed "5 days" for four weeks in a row, because the cause kept moving:
| Week | What the trace showed | Where the right evidence died |
|---|---|---|
| 1 | The 2026 policy lived on the HR portal, which the ingestion job never crawled | Never indexed |
| 2 | Portal added. The 2024 handbook chunk scored 0.814 against the question, the 2026 policy 0.776, and only the top chunk went into the prompt | Retrieved, then ranked out |
| 3 | The top 3 were now kept, and a reranker put the 2026 policy first. The prompt builder sorted chunks by title and stopped at the token budget, so the 2026 policy fell off the end | Lost in assembly |
| 4 | The 2026 policy sat first in the prompt and the 2024 handbook second. The model still answered 5 | Ignored or contradicted by the generator |
The week-2 scores are real cosine similarities from a small open embedding model; you will reproduce them below.
The symptom never changed, and no single prompt tweak would have fixed all four weeks. What fixed them was reading the trace: following the one chunk that holds the answer through every stage and asking where it disappeared. That is the core skill of this lesson. A RAG system is a chain of handoffs, and an answer is only as good as the weakest one.
| # | By the end you can answer |
|---|---|
| 1 | Does this system need retrieval at all? |
| 2 | RAG, one giant prompt, or fine-tuning? |
| 3 | What are the moving parts, and what must stay consistent between them? |
| 4 | Where did this wrong answer come from? |
| 5 | What can I promise the people who rely on it? |
| 6 | Classic, agentic, vectorless or graph? |
| 7 | Did my change help? |
This is the root lesson of the RAG branch. It gives you the map and the diagnostic habit. Classic RAG Pipeline builds the pipeline line by line, and each stage has its own lesson, linked where it comes up. You should already know what an embedding and a similarity score are; if not, start with Foundations of Modern AI Search.
Before any architecture: five questions
Before you draw a single box, pin down what the answer needs. These five questions decide more of the architecture than any library choice:
| Ask | Why it matters |
|---|---|
| Does the model already know this, reliably? | Stable, widely published facts may need no retrieval at all |
| How often does it change? | Facts that change every week cannot live in model weights |
| Who may see which part? | Permissions must be enforced before text reaches the model |
| Must the answer show where it came from? | Citations need an id for every passage, carried through every stage |
| How big is it, and how many questions a day? | Decides whether "put everything in the prompt" is affordable |
For the leave assistant the answers are: no, the policy is private; yearly, with amendments; one chapter (compensation bands) is restricted to HR; yes, HR wants the section cited; and about 180,000 tokens of HR documents, asked about 2,000 times a day. The next two sections turn those answers into a decision.
Why retrieve at all?
The paper that named RAG, Lewis et al. (2020), paired two kinds of memory. Parametric memory is the knowledge stored in a model's weights. Non-parametric memory is an index of documents searched at question time. Almost every trade-off in this branch of the course comes from how differently the two behave:
| Knowledge in the weights | Knowledge you retrieve | |
|---|---|---|
| Changes when | the model is retrained | you re-index a document |
| Covers | what was common in public training data, up to a cutoff | whatever you index, private data included |
| Points to a source | no | yes: document, section, version |
| Scoped per user | no, every user gets the same weights | yes, by filtering before the text reaches the model |
| Fails by | confidently filling gaps | missing, stale or wrong evidence (the rest of this lesson) |
That gives the four reasons to retrieve:
- Freshness. The 2026 policy took effect a month before the question. Re-index it, and the next question sees it.
- Private data. Your policies, tickets, contracts and code were never in any training set, and should not be.
- Provenance. "[Leave Policy 2026, Β§2]" lets a reader check the claim in ten seconds. Weights cannot tell you where a sentence came from.
- Access control. A contractor and an HR partner can ask the same question and be shown different documents, because retrieval can filter per user. Weights cannot forget a chapter for one user.
Why the plain model guessed. Kalai et al. (2025) argue that models hallucinate largely because training and evaluation reward a confident guess over "I don't know"; whether hallucination can be removed entirely is still debated. Either way, retrieval changes what the model guesses from, not the pull toward guessing. When the evidence is missing, a RAG system guesses too, unless you give it an explicit way to say "not in the sources" and check that it uses it.
What retrieval cannot fix. It cannot repair a wrong source: index the 2024 handbook as if it were current, and RAG will cite it faithfully. It cannot make a model reason correctly over the right text. And a top-k search returns a sample of the corpus, not a census. "List every policy that mentions remote work" needs every matching document, and top-5 retrieval will quietly return five.
Your turn. Which of these requests needs retrieval over documents?
- "What is the boiling point of water at sea level?"
- "What did the March 2026 board meeting decide about hiring?"
- "Rewrite this email so it sounds friendlier." (The email is pasted in.)
- "What is my current account balance?"
Check your answer
- No. It is a stable, widely published fact. Retrieval adds cost and a new way to fail.
- Yes. It is private and recent.
- No. Everything the model needs is in the request.
- Not over documents. The balance lives in a database and changes by the minute, so the assistant should query the system of record at question time, as a tool call. An indexed text copy of balances is stale the moment it is built.
Three ways to give a model knowledge
In 2026 you have three serious options, plus a fourth for live data:
- Retrieve (RAG): put the few most relevant chunks in the prompt at question time.
- Long context: send the whole corpus with every question. Several API models now accept about a million tokens; Anthropic's context-window documentation, for example, lists a 1M-token window for most of its current Claude models (checked September 2026).
- Fine-tune: train the knowledge into the weights.
- Call a tool: for live, structured data (balances, stock levels, ticket status), query the system of record at question time.
Long context: simple, and billed per token per question
For a corpus that fits, long context is the simplest thing that could work. There is nothing to chunk, index or rank, and nothing can be "ranked out". It has two costs.
You pay for the whole corpus on every question. Prompt caching softens this: providers bill a reused prompt prefix at a discount. Anthropic, for example, bills cache reads at 10% of the normal input price or less, depending on the model, with a 5-minute default lifetime (prompt-caching docs, September 2026). Here is the leave assistant's input bill:
PRICE = 3.00 # USD per million input tokens: an assumption, check your provider
CACHE_READ = 0.10 # cached tokens billed at 10% of PRICE
HANDBOOK = 180_000 # tokens: the whole corpus in every prompt
OVERHEAD = 600 # system prompt + question
RAG_CONTEXT = 6 * 500 # six retrieved chunks of about 500 tokens
def usd_per_1000_queries(tokens_full, tokens_cached=0):
per_query = (tokens_full * PRICE + tokens_cached * PRICE * CACHE_READ) / 1e6
return 1000 * per_query
print(f"long context, no cache : {usd_per_1000_queries(HANDBOOK + OVERHEAD):7.2f}")
print(f"long context, cached : {usd_per_1000_queries(OVERHEAD, tokens_cached=HANDBOOK):7.2f}")
print(f"RAG, 6 chunks : {usd_per_1000_queries(RAG_CONTEXT + OVERHEAD):7.2f}")
# long context, no cache : 541.80
# long context, cached : 55.80
# RAG, 6 chunks : 10.80
Per 1,000 questions, sending the handbook uncached costs about 50 times as much as RAG's 3,600-token prompts; caching brings that down to about 5 times. At 2,000 questions a day that is roughly USD 1,080, USD 110 and USD 22 a day. Output tokens cost the same in every design, so they are left out. Cache writes are left out too: they cost extra (1.25 times the input price for Anthropic's 5-minute cache) each time the cache has to be rebuilt.
That last point is why the cached line holds only at volume. At 2,000 questions over an 8-hour day, a question arrives about every 15 seconds and a 5-minute cache never goes cold. At 40 a day, one arrives about every 12 minutes, so two out of three find the cache expired and pay the write surcharge instead of the read discount, and the saving mostly disappears. A longer-lived cache changes the sum (Anthropic also sells a 1-hour cache, whose writes cost 2 times the input price), so price your real arrival pattern before you count on the discount.
Longer inputs are read less reliably. Liu et al. (2023) found that accuracy on multi-document question answering was often highest when the relevant passage sat at the start or end of the input, and dropped when it sat in the middle. RULER (Hsieh et al., 2024) tested 17 models that all claimed windows of 32K tokens or more, and found only half of them performed satisfactorily at 32K. Chroma's Context Rot report (2025) measured performance falling with input length across 18 models, even on simple tasks. How large these effects are varies a lot by model and task. Treat them as a reason to measure accuracy at your prompt length, not as a law.
What the head-to-head studies found. Li et al. (2024) compared the two approaches directly. With enough resources, long context beat RAG on average, and RAG cost far less. Their Self-Route method gives the model the retrieved chunks first, lets it judge whether they are enough to answer, and falls back to the full long context only when they are not. It kept performance comparable to long context at much lower cost. The useful question is not "which one wins?" but "which questions need which?"
Fine-tuning: behavior, not a changing fact base
Fine-tuning changes weights, so it inherits every weakness in the left column of the table above: no source to cite, no per-user scoping, and a new training run every time the policy changes. It is also a weak way to add facts. Ovadia et al. (2023) found that RAG consistently outperformed unsupervised fine-tuning for injecting knowledge, both for facts seen in pretraining and for new ones.
Fine-tuning earns its cost for behavior: an output format, a tone, domain jargon, or skill at using retrieved documents. RAFT (Zhang et al., 2024) fine-tunes a model to answer from retrieved documents, ignore the distractors among them, and quote the passage it relied on. That is fine-tuning and RAG working together, not competing.
Choosing
| Requirement | Leans toward | Because |
|---|---|---|
| Corpus fits the window with room to spare; few questions a day | Long context | Nothing to chunk, index or rank, and the bill stays small at low volume even without a cache |
| Questions need the whole corpus ("list everyβ¦") and it fits | Long context | Top-k returns a sample |
| Corpus far larger than any window | RAG | It cannot be sent at all |
| Users may see different documents | RAG with server-side filters | One shared prompt cannot hide a chapter from one user |
| High volume, tight cost per question | RAG | Thousands of tokens per question instead of the whole corpus |
| Answers must cite the passage | RAG, or long context with passage ids in the prompt | Citations need addressable passages |
| Format, tone, jargon, habits | Fine-tuning | Behavior, not facts |
| Live structured data | Tool call or query | A snapshot goes stale |
Your turn: change one requirement. The leave assistant (180,000 tokens, 2,000 questions a day) could run on long context today if you accept the bill. What changes in each case?
- Legal and IT documents join the corpus, which grows to 40 million tokens.
- Contractors must never see the 12-page compensation-bands chapter.
- Volume drops to 50 questions a day, and most questions are "what changed between the old and the new policy?"
Check your answer
- RAG. Forty million tokens is 40 times a one-million-token window. It cannot be sent at all, cached or not.
- Either can work, with care. Long context can serve two cached prompts, one per permission group. That holds for two groups and breaks when permissions are set per document or per team. RAG enforces the rule with a filter on each chunk's permission field before anything is returned. What works in neither design is an instruction such as "do not reveal compensation data to contractors": by then the text is already in the prompt, and an instruction is not an access control.
- Long context becomes attractive. At 50 questions a day even the uncached bill is about USD 27 a day (0.5418 Γ 50), and comparing two policies needs both of them in full, which chunk retrieval may split. Measure accuracy at that prompt length before you commit.
The map: six stages, one chain of custody
RAG is a pattern, not a product: retrieve evidence for a question, augment the prompt with it, generate an answer from it. The evidence can come from a vector index, a keyword index, a SQL database, a knowledge graph or a web search API. What makes a system RAG is that shape, not the store. A production system wraps the shape in three more stages: two that run before any question arrives (ingest and index) and one that watches everything (evaluate).
OFFLINE: runs whenever a source changes
sources ββββββΊ 1 INGEST βββββββββββΊ 2 INDEX ββββββββββββΊ index
wiki, PDFs, fetch, parse, chunk, embed, vectors, terms
tickets, DB dedupe, attach store with + payload: doc_id,
version + ACL payload version, ACL
β
ONLINE: runs for every question β search
βΌ
question βββββββββββββββββββββββββββββββββββββββββββββΊ 3 RETRIEVE
rewrite, search,
filter by user, rerank
β candidate pool
βΌ
4 AUGMENT
keep the top chunks,
fit them to a token
budget, order, label
sources
β prompt holding
β the context set
βΌ
5 GENERATE βββββββΊ answer
answer from the + cited ids
sources, cite them,
or abstain
6 EVALUATE βββββ trace of every stage: pool, selected ids, prompt, answer
labeled questions, one number per stage
| Stage | Job | Decisions that live here | Go deeper |
|---|---|---|---|
| 1 Ingest | Get every source in, parsed, with identity, version and permissions | Connectors, parsing PDFs and tables, change detection, deletes | Data Pipeline & Indexing, Document Processing, Data Freshness & Lifecycle |
| 2 Index | Turn documents into searchable units | Chunk size and boundaries, embedding model, lexical index, payload fields | Smart Chunking, Embedding Pipeline, Vector Database Architecture |
| 3 Retrieve | Find a candidate pool for this question and this user | Query rewriting, lexical + dense search, filters, reranker, pool size | Query Understanding & Intent, Hybrid Retrieval Systems, Metadata & Filtering, Reranking Models |
| 4 Augment | Turn the top candidates into the context set: the chunks that actually go into the prompt | The cut-off (how many chunks to keep), token budget, order, source labels, marking retrieved text as untrusted | Context Augmentation |
| 5 Generate | Answer from the context, cite it, or abstain | Grounding instruction, citation format, when to say "not in the sources" | Classic RAG Pipeline |
| 6 Evaluate | Know whether each stage works | Labeled questions, a metric per stage, latency and cost | Evaluation & Quality Metrics |
π§ A chain of custody. In court, evidence is worthless if nobody can show where it was at every moment between the scene and the jury. Treat a chunk the same way: from ingest to citation you should be able to say where it was, and your trace log is the custody record. The analogy stops in one place. Court evidence must arrive unchanged, while a chunk is deliberately transformed at almost every stage (split, embedded, truncated, reordered). So the trace must record what form reached each stage, not only that something arrived.
Three things must hold across every handoff:
- The same representation on both sides. The query must be embedded by the same model, with the same prefix or instruction, that embedded the chunks. Some models make this explicit: the E5 model card says to start queries with
query:and passages withpassage:, or expect worse results. - Identity travels with the text. Every chunk carries its document id, chunk id, version or effective date, source link and permissions from ingest to citation. Without them you cannot filter, cite, delete or debug.
- Every handoff is logged. For each question, log the candidate pool (ids and scores), the ids selected for the prompt, the exact prompt text, the answer and the ids it cites. Without that log, the four-week story is four weeks of guessing.
Predict first. A team upgrades its embedding model. New documents are embedded with the new model; old chunks keep their old vectors; every query now uses the new model. What happens on the next query if:
- (a) the old model produced 768-dimensional vectors and the new one produces 1,024?
- (b) both produce 1,024-dimensional vectors?
Check your answer
(a) A loud failure. The collection was created for 768 dimensions, so every insert and every query of the wrong size is rejected. Qdrant 1.19 returns 400 Bad Request with "Vector dimension error: expected dim: 768" and the size it got; FAISS raises an error. Annoying, but you find it in minutes.
(b) A silent failure. The sizes match, so every call succeeds. New chunks are compared with the query in the space they were embedded in. Old chunks are compared across two different spaces. For unrelated models their scores are noise that still looks like a similarity, and they drift out of the results without a single error; for closely related models the two spaces can partly agree, so results degrade rather than collapse, which is even harder to notice. The fix is to re-embed everything with the new model, usually into a new collection before switching over. A model change is a re-index; Vector Embeddings covers it in depth.
Where the evidence died: a failure taxonomy
Take a question whose right answer you know, and the chunk that contains it: the gold chunk. Walk it through the trace, checkpoint by checkpoint. Three sets on the trace are easy to blur, so pin them down once:
- the candidate pool is what retrieval returns, ranked (often 50 to 200 chunks);
- the selected chunks are the top of that ranking, up to your cut-off (say the top 5), which the prompt builder is asked to include;
- the context set is what actually ended up in the prompt.
The first checkpoint the gold chunk fails is where the symptom starts. Everything after it never saw the chunk, so tuning it cannot fix that question. The fix can still live earlier: week 2's was version metadata written at ingest.
| Bucket | Test on the trace | First move | What the move costs |
|---|---|---|---|
| Not indexed | The index does not match the source: the gold document is missing, or its indexed text or metadata (path, permissions, status) is older than the source | Coverage and freshness checks at ingest | Pipeline work, re-index compute |
| Not retrieved | In the index, but not in the candidate pool | Hybrid retrieval, consistent embedding, re-chunking | A lexical index to run and sync; re-embedding |
| Ranked out | In the pool, but ranked below the cut-off, so never selected | Version metadata, deduplication, rerank a bigger pool, a larger cut-off | Schema work; reranker latency; more prompt tokens |
| Lost in assembly | Selected, but missing from the prompt, or cut | Fill the budget in rank order; log the prompt | Almost nothing: most are bugs |
| Ignored or contradicted | In the prompt, but the answer does not use it or contradicts it | Remove conflicts upstream, fewer and better chunks, cite-or-abstain, faithfulness checks | Prompt work; judge calls to check answers |
If a question has no gold chunk anywhere, because the answer is not in your documents, there is one correct outcome: the system says it cannot answer from its sources. Put such questions in your test set too. The subsections below give each bucket's usual causes and how to confirm them.
Not indexed: the index does not match the source
Detection needs no machine learning. Look up the gold document by id in the index and compare it with the source: its text, and the metadata that retrieval filters on (path, permissions, status). If it is missing, or any of that is older than the source, stop here: retrieval is working exactly as designed on the wrong corpus.
Typical causes are a source the connector does not crawl (week 1), a scanned PDF with no text layer that parsed to an empty string, a table the parser dropped, and a nightly sync for a policy that changed at 9:00. The quiet one is a document deleted at the source whose chunks are still in the index and still being cited. It shows up on the trace in a later bucket, because the dead chunks outrank or contradict the live ones, so the check has to happen at ingest. The moves are unglamorous: a coverage report (documents at the source vs documents in the index), change detection, delete propagation and alerts on parse failures. They live in Data Pipeline & Indexing and Data Freshness & Lifecycle.
Not retrieved: indexed, but the search never found it
The gold chunk is in the index but not in the candidate pool. The usual suspects, each with a check:
- An exact identifier the dense model blurs. Part numbers, error codes and clause ids (
XK-4471,ECONNREFUSED) are weak spots for embeddings. Check: does a keyword search find the chunk? If it does, add lexical retrieval β Hybrid Retrieval Systems, Sparse vs Dense Retrieval. - Different representations on the two sides: another model, a missing
query:prefix, a different normalization. Check: embed the query exactly as the chunks were embedded. - The answer sits past the model's input limit.
all-MiniLM-L6-v2reads at most 256 word pieces and silently ignores the rest, so the end of a 600-word chunk never reaches its vector. Check the chunk's length in tokens β Smart Chunking. - The question uses other words. "Roll over my PTO" against "carry over vacation days" can defeat keyword search, and sometimes dense search too. Check a rewritten query β Query Understanding & Intent.
- A filter excluded it: a filter that is too strict, or a permission or status field that was never set. Check: run the same query without filters.
- The approximate index missed it. Check: re-run the query with exact (brute-force) search. If exact search finds the chunk, tune the index (ANN Algorithms). If it does not, the problem is the representation, and no index setting will fix it.
Ranked out: found, then beaten
The gold chunk is in the candidate pool but ranked below the cut-off, so it is never selected. Week 2 is the classic case. Run it yourself (the model downloads on first use):
import numpy as np
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2") # 384 dims
chunks = [ # (chunk id, status, text)
("hb24-4.2", "superseded",
"Employee Handbook 2024, section 4.2: Unused vacation days may be carried over to the "
"next calendar year, up to a maximum of 5 days. Days above the cap are forfeited on January 1."),
("pol26-carry", "current",
"Leave Policy 2026 (effective January 1, 2026): Employees may carry over up to 10 unused "
"vacation days into the next year. Submit a carry-over request in the HR portal by January 31."),
("faq-7", "current",
"Leave FAQ, question 7: Can I carry over vacation time? Yes. See the current leave policy "
"for the cap and the request deadline."),
("hol-26", "current",
"Office closures 2026: the office is closed on 12 public holidays. Public holidays do not "
"count against your vacation balance."),
("exp-3", "current",
"Expense policy, section 3: Carry receipts for every business expense and submit them "
"within 30 days of the trip."),
("par-1", "current",
"Parental leave: 16 weeks paid leave for the primary caregiver, 6 weeks for the secondary caregiver."),
]
vecs = model.encode([text for _, _, text in chunks], normalize_embeddings=True)
query = "How many unused vacation days can I carry over into next year?"
q = model.encode(query, normalize_embeddings=True)
scores = vecs @ q # unit vectors, so this is cosine similarity
for rank, i in enumerate(np.argsort(-scores), start=1):
chunk_id, status, _ = chunks[i]
print(f"{rank}. {chunk_id:<12} {status:<11} {scores[i]:.3f}")
Predict first: which chunk ranks first, the superseded 2024 handbook or the current 2026 policy?
Check your answer
1. hb24-4.2 superseded 0.814
2. pol26-carry current 0.776
3. faq-7 current 0.574
4. hol-26 current 0.501
5. exp-3 current 0.472
6. par-1 current 0.316
The superseded handbook wins by 0.038. Both chunks talk about carrying unused vacation days into the next year. An embedding encodes what a text is about; which version is in force is invisible to similarity. With a cut-off of one chunk, only the handbook is selected, and the answer is "5 days". The expense chunk, which shares only the words "carry" and "days" with the question, ranks fifth.
The moves, cheapest first:
- Tell the index which version is current. Store
statusoreffective_fromwith every chunk at ingest, and filter or boost on it at query time β Metadata & Filtering. Often the cheapest fix is not indexing superseded documents at all. - Deduplicate. Five copies of the same paragraph (drafts, exports, email attachments) can fill a top 5 on their own.
- Rerank a bigger pool. A cross-encoder reads the question and each chunk together and orders a candidate pool (commonly 50 to 200 chunks) better than first-stage similarity does. The cost is one model call per pair, and it can only reorder what the first stage returned: a chunk that never reached the pool is a not-retrieved problem β Reranking Models.
- Raise the cut-off when questions have several parts, as long as the selected chunks still fit the token budget (next section).
β οΈ A reranker would not reliably have fixed week 2. It judges relevance, and the 2024 handbook is relevant to the question. Only metadata knows it is out of date.
Lost in assembly: selected, then dropped on the way into the prompt
The gold chunk made the cut, but it never reaches the prompt, or it reaches it cut. Nothing here is a model problem. It is ordinary code between the retriever and the LLM. Here is week 3's prompt builder:
ranked = [ # after reranking, best first
{"id": "pol26-carry", "title": "Leave Policy 2026", "tokens": 620},
{"id": "hb24-4.2", "title": "Employee Handbook 2024", "tokens": 540},
{"id": "faq-7", "title": "Leave FAQ", "tokens": 480},
]
BUDGET = 1_200 # context tokens left after system prompt, question and answer reserve
def assemble(chunks, budget):
chunks = sorted(chunks, key=lambda c: c["title"]) # "so it reads like a document"
picked, used = [], 0
for c in chunks:
if used + c["tokens"] > budget:
break
picked.append(c["id"])
used += c["tokens"]
return picked
print(assemble(ranked, BUDGET))
What does this print?
Check your answer
['hb24-4.2', 'faq-7']. Sorting by title puts "Employee Handbook 2024" (540 tokens) and "Leave FAQ" (480) first, using 1,020 tokens. The 620-token 2026 policy would bring the total to 1,640, over the 1,200 budget, so the loop stops. The reranker's first choice is gone, and the superseded handbook made it in.
Two fixes: choose in rank order, and skip a chunk that does not fit instead of stopping. Reorder only what you chose, if you must.
def assemble(chunks, budget):
picked, used = [], 0
for c in chunks: # chunks arrive best first
if used + c["tokens"] <= budget: # skip what does not fit, keep going
picked.append(c["id"])
used += c["tokens"]
return picked
print(assemble(ranked, BUDGET)) # ['pol26-carry', 'hb24-4.2']
That uses 1,160 tokens. The FAQ (480 more) no longer fits, and skipping it is a budget decision, not a bug: three selected chunks were more than a 1,200-token budget can hold. Log such skips, and if they start dropping gold chunks, lower the cut-off, raise the budget or split long chunks. Notice what is still in the prompt: the superseded handbook. That is week 4.
Other assembly bugs look similar: cutting a chunk by characters when the sentence with the number came last; deduplication that keeps the older of two near-identical chunks; a template that silently drops chunks whose text field became empty after a schema change; a framework that trims an oversized prompt without telling you. The test is always the same: log the exact prompt text and check that the gold chunk's text is in it. Ordering and budgets in depth β Context Augmentation.
Ignored or contradicted: the model had it and did not use it
The gold chunk is in the prompt, and the answer still disagrees with it. In week 4 the 2026 policy came first, the 2024 handbook second, and the answer was 5.
Why models do this is an active research question, and the answer depends on the model and the evidence. Longpre et al. (2021) found that question-answering models over-relied on memorized knowledge when a passage contradicted it. Xie et al. (2024) found that large language models were often receptive to coherent evidence that contradicted their memory, but showed confirmation bias when the context held evidence on both sides, which is exactly what week 4's prompt did. Other causes are more mundane: a weak or missing instruction to answer only from the sources, and long prompts full of near-misses in which the relevant passage sits in the middle (Liu et al., 2023; the size of the effect varies by model and task).
The moves:
- Remove the conflict upstream. Filtering out superseded versions, the ranked-out fix, also fixes week 4.
- Label every chunk with its source and date, so the model can prefer the newer one and cite it.
- Require a citation for each claim, and an explicit "not in the sources" answer. Templates β Classic RAG Pipeline.
- Send fewer, better chunks. Every extra near-miss is a distractor.
- Check the answer against the context. A faithfulness check (a person or an LLM judge asks whether each claim is supported by the chunks in the prompt; whether it is supported by the chunk it cites is the stricter citation check) catches this bucket, and nothing upstream will β Faithfulness Testing.
β οΈ Lowering the temperature makes answers more repeatable. It does not make the model prefer your context: if "5" is the model's most likely answer, greedy decoding (temperature 0, where the API offers it) will keep saying "5".
Put it in code, then count
Once every question in your test set has a gold chunk id and every run is logged, locating the failure takes a few lines, and the histogram tells you what to work on next. answer_ok means the answer is correct and supported by the gold chunk, as judged by a person or by an LLM judge you have checked against people.
from collections import Counter
BUCKETS = ["not indexed", "not retrieved", "ranked out", "lost in assembly",
"ignored or contradicted", "ok"]
def locate_failure(gold, indexed_current, pool, selected, prompt_ids, answer_ok):
"""First checkpoint at which the gold chunk went missing."""
if not indexed_current: # missing, or indexed text/metadata older than the source
return "not indexed"
if gold not in pool:
return "not retrieved"
if gold not in selected:
return "ranked out"
if gold not in prompt_ids:
return "lost in assembly"
if not answer_ok:
return "ignored or contradicted"
return "ok"
G = "pol26-carry"
# gold, indexed and current?, candidate pool (top ids), selected, ids found in the prompt, answer ok?
traces = [
(G, False, ["hb24-4.2", "faq-7", "hol-26"], ["hb24-4.2"], ["hb24-4.2"], False), # week 1
(G, True, ["hb24-4.2", G, "faq-7"], ["hb24-4.2"], ["hb24-4.2"], False), # week 2
(G, True, [G, "hb24-4.2", "faq-7"], [G, "hb24-4.2", "faq-7"],
["hb24-4.2", "faq-7"], False), # week 3
(G, True, [G, "hb24-4.2", "faq-7"], [G, "hb24-4.2", "faq-7"],
[G, "hb24-4.2"], False), # week 4
("sick-2", True, ["sick-2", "par-1"], ["sick-2"], ["sick-2"], True),
("vpn-9", True, ["vpn-2", "vpn-4"], ["vpn-2"], ["vpn-2"], False),
("tax-1", False, [], [], [], False),
("par-1", True, ["par-1", "sick-2"], ["par-1"], ["par-1"], True),
]
counts = Counter(locate_failure(*t) for t in traces)
for bucket in BUCKETS:
print(f"{bucket:<24} {counts[bucket]}")
# not indexed 2
# not retrieved 1
# ranked out 1
# lost in assembly 1
# ignored or contradicted 1
# ok 2
Read the histogram, not individual failures. Here two of eight questions died before retrieval even ran, so no reranker or prompt change could have saved them. On a real system, run this over 50 to 100 labeled questions after every change.
Your turn: classify by hand. For each trace, name the bucket and the first move.
- The gold chunk
sla-3is 14th in a reranked 100-chunk candidate pool. The top 5 are selected for the prompt. - The answer says the warranty lasts one year. The gold chunk in the prompt says two years, and no other chunk in the prompt mentions the warranty.
- The gold document was moved to a new folder at the source last week. The index still holds it under its old path, and a retrieval filter
path starts with /policies/current/excludes it.
Check your answer
- Ranked out. The pool is already reranked, so check why five chunks beat it (duplicates? stale versions?) and remove what should not compete: deduplicate, or filter out superseded versions. Also check whether the question needs more than five chunks.
- Ignored or contradicted. Nothing in the prompt says one year, so the number came from the model's prior. Tighten the cite-or-abstain instruction and add a faithfulness check. This one is worth reporting to whoever picks the model.
- Not indexed, in a form that is easy to miss: the index holds the document, but its path is older than the source's. If you only check that the id exists, the trace looks like not retrieved, because the filter keeps the chunk out of the pool. That is why the ingest check compares metadata with the source, not just text. First move: make ingest propagate moves (and permission changes) to the payload, then re-sync this document.
What RAG promises, and what it doesn't
| RAG gives you | RAG does not give you |
|---|---|
| Fewer unsupported answers, when retrieval finds the right evidence | The end of hallucination: with missing or conflicting evidence the model can still guess |
| A record of what the model was shown, if you log the prompt | Proof that each sentence is supported: citations can be wrong or decorative |
| Answers as fresh as your last sync | Freshness beyond your last sync |
| Per-user access control, if enforced as a server-side filter during retrieval | Access control through instructions in the prompt |
| Evidence you can inspect | Safety from what that evidence says: retrieved text can carry instructions |
Four of those rows deserve a sentence each.
Retrieval caps the grounded answer, not the answer. When the evidence never arrives, the model can still answer from its weights, sometimes correctly. Suppose 200 labeled questions show 176 correct answers, but the gold chunk reached the context set for only 140 of them. Then at least 36 correct answers (18% of all questions) came without the evidence. End-to-end accuracy looks fine while retrieval is broken, which is why you measure retrieval on its own.
A citation is the model's claim about its source, not proof. The model writes it like any other text, so it can point to a chunk that says something else, or to one that was never in the prompt. Check the cited id against the logged context set, then check that the claim follows from that chunk. The second check is a support test, sentence by sentence or claim by claim β Citation Coverage, Grounding & Hallucination Control.
Access control happens before the prompt. Filter by the user's permissions inside the retrieval query, on the server, and deny by default. A line in the system prompt saying "do not reveal X" is a request, not a control β Metadata & Filtering, Infrastructure & Security.
Retrieved text is untrusted input. A document can contain text aimed at the model ("ignore your previous instructions andβ¦"). Greshake et al. (2023) showed such indirect prompt injection working against real LLM-integrated applications. How to delimit and handle it β Context Augmentation.
The family: classic, agentic, vectorless, graph
All four variants keep the retrieve β augment β generate shape. They differ in what they retrieve from and in who decides what to retrieve.
| Variant | What changes | Fits questions like | Main cost | Lesson |
|---|---|---|---|---|
| Classic | One retrieval pass, one generation | "What is the carry-over cap?" | Fails quietly on questions that need several lookups | Classic RAG Pipeline |
| Agentic | The model calls retrieval as a tool and decides whether, where and how often to search, and when to stop | "Compare our 2024 and 2026 leave rules for part-timers" | Several LLM calls per answer: latency, cost, harder testing | Agentic RAG Systems |
| Vectorless | Retrieval without embeddings: keyword search, SQL over structured data, or an LLM navigating a document's table of contents (for example PageIndex) | "How many P1 tickets are open for the Berlin team?" | Needs structure or exact vocabulary; generating SQL or walking a document tree adds LLM calls and their own failure modes | Vectorless RAG |
| Graph | A graph of entities and relationships extracted from the corpus, sometimes with summaries of related groups | "What themes recur across 2,000 incident reports?" | LLM calls at indexing (the original GraphRAG) or at query time (lighter variants such as LazyGraphRAG), entity resolution, keeping the graph in sync | Graph Rag |
The graph row's example question is modelled on the original GraphRAG paper, Edge et al. (2024), which targeted "global" questions about a whole corpus, such as its main themes, that no top-k set of chunks can answer.
Where agentic latency comes from. Assume one LLM call takes 1.5 s and one retrieval takes 0.1 s. A classic pipeline costs 0.1 + 1.5 = 1.6 s. An agent that plans (one call), searches and reads three times (3 Γ (0.1 + 1.5) s), then writes the answer (one call) takes 1.5 + 4.8 + 1.5 = 7.8 s, and only 0.3 s of that is retrieval. The extra time is LLM calls, not searching. Running independent searches in parallel and capping the number of steps is how you win it back β Agentic RAG Systems, and for splitting a question into sub-questions, Query Decomposition.
Start classic, then route the exceptions. Most systems begin as one classic pipeline. The failure histogram shows which question types it fails, and those types get routed to an agent, a SQL tool or a graph. Li et al.'s Self-Route, earlier in this lesson, is the same idea applied to RAG versus long context.
Your turn. Procurement asks: "Which of our 40 supplier contracts renew in Q4 and contain an auto-renewal clause?" The contracts are PDFs of 30 to 60 pages each. A teammate proposes classic RAG with the top 8 chunks. What goes wrong, and what would you build?
Check your answer
Top 8 returns eight chunks. The question needs a renewal date and an auto-renewal flag for all 40 contracts, so it is an aggregation over the whole set, and a top-k sample will quietly leave contracts out. Build it as structured data instead. At ingest, extract renewal_date, auto_renewal and the page each value came from for every contract, with an LLM, and check a hand-labeled sample of about 10 contracts. Store the fields in a table and answer with a query: that is one vectorless route. Keep chunk retrieval for questions like "what exactly does clause 14 of the Acme contract say?". An agent that reads all 40 contracts also works, at 40 or more LLM calls per question, which is acceptable only if the question is rare.
Measure it before you tune it
The four-week story took four weeks because nobody had a test set. The minimum that turns RAG work into measurement:
- 50 to 100 real questions from logs or users, each labeled with its gold chunk or document ids and a short reference answer. Include some whose answer is not in the corpus, to test abstention.
- A trace for every run, as above.
- One number per stage, plus the two that pay the bills:
| Stage | What to measure |
|---|---|
| Ingest | Coverage: share of gold documents present in their current version |
| Retrieve | Hit rate and recall at the pool size (say @100) and at the cut-off (say @5) |
| Augment | Gold-in-prompt rate: share of questions whose gold chunk text appears in the logged prompt; compare it with recall at the cut-off (a gap is a bug, or a cut-off larger than the budget can hold) |
| Generate | Correct and supported answers; correct abstentions on unanswerable questions |
| Whole system | p95 latency; cost per 1,000 questions |
Two retrieval numbers are easy to confuse. Hit rate@k (other lessons write hit@k) is the share of questions with at least one relevant chunk in the top k. Recall@k is, for each question, the relevant chunks in the top k divided by all relevant chunks for that question, averaged over questions. Hit rate is never lower than recall, and the two agree only when every question finds either all of its relevant chunks or none:
| Question | Relevant chunks | Found in top 5 | Hit? | Recall@5 |
|---|---|---|---|---|
| q1 | 1 | 1 | yes | 1.0 |
| q2 | 2 | 1 | yes | 0.5 |
| q3 | 3 | 3 | yes | 1.0 |
| q4 | 1 | 0 | no | 0.0 |
Hit rate@5 = 3/4 = 0.75. Mean recall@5 = (1 + 0.5 + 1 + 0) / 4 = 0.625. The whole gap comes from q2, which found one of its two chunks; q3 has three relevant chunks but found all of them, so it scores the same on both. Questions with several parts need recall, because one hit is not enough to answer them. MRR, nDCG and friends β Retrieval Metrics. Judging answers β Generation Quality.
Then change one thing at a time and re-run. With a test set, each of the four weeks is a one-day fix: the histogram names the bucket, and the re-run shows whether the fix worked and what it cost in latency and money.
Final round: no label on the failure
Real bug reports do not say which bucket they belong to. Read each trace, name the first checkpoint the gold chunk failed, and pick the first move.
Case 1: the confident refund
A retail assistant says returns are accepted "within 14 days". The gold chunk returns-2 ("30 days from delivery") is first in the candidate pool, first among the selected chunks, and present in the logged prompt. The second chunk in the prompt is from a company blog post: "Most retailers accept returns within 14 days."
Check your answer
Ignored or contradicted, with a root cause upstream. The model went with the blog's number, a conflict the prompt handed it. Keep blog posts out of the policy index, or mark source authority in the payload and filter on it. Then add a faithfulness check, because the next conflict will come from somewhere else.
Case 2: the error code nobody can find
"What does error E-7031 mean?" The candidate pool holds 100 chunks from dense search only, and none of them contains E-7031. A keyword search over the same corpus returns errors-e7xxx, the chunk that defines E-7031, first.
Check your answer
Not retrieved. The chunk is indexed (keyword search finds it), and the dense model blurs the exact code. Add lexical retrieval and fuse the two lists. The cost is a second index to run and keep in sync.
Case 3: the half answer
"What are the notice periods for employees and for contractors?" The five-chunk context set holds four near-identical copies of the employee notice-period paragraph (from the handbook, a PDF export, the FAQ and an old email) plus an onboarding chunk. The contractor paragraph is 9th in the candidate pool.
Check your answer
Ranked out, because duplicates crowd the top 5. Deduplicate at ingest, or diversify the results so near-copies do not fill the context set. A two-part question may also deserve a bigger context set, or splitting into two sub-questions.
Case 4: the policy that came back from the dead
A travel policy was deleted from the wiki three weeks ago and replaced by a new one. The assistant still quotes the old hotel limit. The trace shows the old policy's chunk first in the candidate pool and the new policy's chunk 7th; the top 5 are selected for the prompt.
Check your answer
Ranked out on the trace, but the root cause is at ingest: the index no longer matches the source. Deletes are not propagated, so the dead document's chunks still compete with the new policy, and a reranker would happily keep ranking them, because they are relevant. Add delete propagation and a periodic reconciliation between source and index, then remove the orphaned chunks.
Cheat sheet: symptom β stage β move β cost
| Symptom | Likely stage | First move | Cost |
|---|---|---|---|
| Cites an old version or a deleted document | Ingest: sync | Change detection, delete propagation, version metadata | Pipeline work |
| Exact code or part number never found | Retrieve: dense only | Add lexical retrieval and fuse | A second index to sync |
| Worse after a model upgrade, no errors | Index: mixed embedding spaces | Re-embed everything with one model | A full re-embed |
| Right chunk ranked just below the cut-off, top slots full of copies | Retrieve: ranking | Deduplicate, rerank a bigger pool | Reranker latency |
| Right chunk selected, missing from the prompt | Augment: usually a bug | Fill the budget in rank order; log the prompt | None |
| Right chunk in the prompt, answer contradicts it | Generate | Remove conflicting sources, cite-or-abstain, faithfulness check | Judge calls |
| "List everyβ¦" answers are incomplete | Wrong variant | Structured extraction and a query, or long context if it fits | An extraction pipeline, or tokens |
| Users see documents they should not | Retrieve: filters | Server-side permission filter, deny by default | Permission sync |
Before moving on, take one wrong answer from a system you know, or week 4 of the story, and explain it aloud from a blank page: which chunk is the gold chunk, the first checkpoint it failed, why nothing after that point could recover, the cheapest move that fixes it, what that move costs, and which number on your labeled set would prove the fix worked. If you can do that, the rest of the RAG branch is detail on stages you already know how to find.
Next: Classic RAG Pipeline builds the whole pipeline in plain Python, Context Augmentation covers what goes into the prompt and in what order, and Agentic RAG Systems covers the cases where one pass is not enough. Vectorless RAG and Graph Rag cover the other two members of the family.