Appearance
Retrieval-Augmented Generation (RAG)
RAG (Retrieval-Augmented Generation) is the paradigm formally proposed in 2020 by Lewis et al. at Meta (then Facebook AI): before generating, retrieve relevant passages from an external knowledge base, stuff them into the prompt, and let the model take an open-book exam. Six years on, it remains the first cornerstone of enterprise LLM adoption — and the system newcomers most often "stand up in three days and spend three months failing to tune." This page skips the concept primer and dissects every link of a production-grade RAG pipeline, telling you where each step's pitfalls live.
1. What Problem RAG Solves
RAG's reason for being boils down to three hard constraints:
- Knowledge cutoff. The model's parametric knowledge freezes at its training-data cutoff. A 2026 model knows nothing about last week's product docs — unless you feed them at inference time.
- Hallucination. An LLM won't say "I don't know" about things it doesn't know; it fabricates with equal confidence. RAG anchors answers in retrieved text, and paired with a citation mechanism, makes answers checkable.
- Private data. An enterprise's internal docs, codebase, and ticket system will never enter the training set. Fine-tuning can train in "tone," but training massive, frequently updated business knowledge into parameters is neither economical nor realistic — RAG is the right path for knowledge injection; fine-tuning is the path for behavior calibration. Don't confuse the two.
When not to use RAG
If your knowledge base is small (a few hundred pages or fewer) and most questions require reading the whole document, stuffing everything into a long context window is often simpler and more accurate than RAG. RAG's value lies in selecting from corpora too large to fit into context. Using RAG on a corpus smaller than the context window trades one nonexistent problem for retrieval errors.
One more clarification: RAG is not the opposite of an agent — it's one of the agent's most important knowledge-acquisition tools. In the framework of What Is an AI Agent, naive RAG is a fixed pipeline, while Agentic RAG turns retrieval into a tool the agent can call.
2. The Full Pipeline, Dissected
A complete RAG pipeline splits into offline indexing (ingestion) and online query stages:
Offline indexing (triggered by document changes)
Document ──► Parse ──► Chunk ──► Embed ──► Write to index
│ │ │ │
▼ ▼ ▼ ▼
PDF/table chunk vector Vector store +
OCR/structure metadata (dense) BM25 inverted index
Online query (per request)
User question ──► Rewrite ──► Retrieve ──► Rerank ──► Assemble prompt ──► Generate
│ │ │ │
▼ ▼ ▼ ▼
HyDE/ top-k cross-encoder LLM +
multi-query recall fine-ranking top-n citationsNow the engineering details, section by section.
Document parsing: garbage in, garbage out
Parsing is the dirtiest and most underrated link in the chain. Key points:
- PDFs are the disaster zone. Multi-column layouts, headers and footers, tables, figure captions — naive text extraction turns it all into porridge. In practice, use dedicated document-parsing tools (open-source options like Unstructured, Docling, and Marker, or cloud document-intelligence services) to convert PDFs into Markdown/JSON with structure markers, rather than extracting bare text.
- Tables need special handling. Flattening a table into text almost always destroys row-column relationships. Common approaches: convert tables to Markdown tables or HTML tables, or generate a natural-language summary of each table to participate in retrieval.
- Preserve metadata. Source document, section titles, page numbers, update times — these both drive later filtering and anchor answer citations. Metadata dropped at parse time can't be recovered later.
Chunking: it caps your retrieval quality
The goal of chunking: every chunk is semantically self-sufficient, and the granularity matches the retrieval scenario. Full comparison in Section 3.
Embedding: the ticket into semantic space
Embedding models map text to vectors. Selection points:
- Domain match matters more than leaderboard scores. A model topping the general MTEB leaderboard may underperform on your legal/medical/code corpus. The most reliable method: run a small-scale evaluation on your own query-document pairs.
- For multilingual corpora, pick a multilingual model. Mixed Chinese-English (or any multilingual) corpora will suffer noticeably with a purely English-optimized model. Multilingual models like BGE-M3 output both dense and sparse vectors, fitting hybrid retrieval.
- Dimensions and cost. Higher dimensions bring slightly better precision at proportionally higher storage/compute cost. Many models support Matryoshka truncation (the first N dimensions remain usable) — a practical trick for storage-sensitive setups.
Once vectors are written, switching models means rebuilding the entire index — the most expensive decision in the indexing stage. Think before you commit.
Retrieval and reranking
The online stage's core is "two-phase retrieval": first cheaply recall top-k with a cheap method (vectors/BM25; k usually 20–100), then precisely rank top-n (usually 3–10) with an expensive model (a cross-encoder reranker). Details in Section 4.
Generation: even the last step can waste everything
- Context assembly. Concatenate retrieval results in reranked order, with source metadata attached. Watch "lost in the middle" — models use the middle of the context less well than the ends, so put key evidence at either end.
- Explicit constraints. Instruct the model to "answer only from the given material, and state plainly when the material is insufficient," and require citations. This is the highest-ROI single step for pushing hallucination down.
- Context budget. More retrieval results is not better. Irrelevant chunks dilute attention and inflate cost — that's context engineering's ledger to keep.
3. Chunking Strategies Compared
| Strategy | Approach | Pros | Cons | Suited scenarios |
|---|---|---|---|---|
| Fixed-length chunking | Hard-cut by token count with overlap (e.g. 512 tokens + 15% overlap) | Simplest to implement; predictable behavior | Cuts sentences/paragraphs; broken semantics | Fast prototypes; a fallback for messy corpora |
| Recursive / structure-aware chunking | Cut on natural boundaries (headings, paragraphs, sentences) first; degrade to hard cuts only for overlong pieces | Preserves document structure; chunks semantically complete | Depends on the document being well structured | The default for most Markdown/technical docs |
| Semantic chunking | Cut at abrupt changes in embedding similarity between adjacent sentences | Boundaries align with semantic shifts | Expensive; unstable boundaries; hard to debug | Long unstructured continuous text (transcripts, etc.) |
| Parent-child chunks | Embed and retrieve with small chunks; return the owning large chunk to the LLM on a hit | Both retrieval precision and generation context | More complex index; maintain a two-level mapping | Answers scattered locally but needing context to understand |
| Contextually enriched chunks | After chunking, have an LLM write a short positioning blurb for each chunk, prepended before embedding | Markedly mitigates "chunks lose meaning out of context" | One extra LLM call per chunk at indexing; higher cost | High-bar enterprise knowledge bases |
Several field-tested judgments:
- Default to "structure-aware + overlap." Recursive chunking, as represented by LangChain's
RecursiveCharacterTextSplitter, is the right starting point for most projects. Don't reach for semantic chunking first. - There's no silver bullet for chunk size. The empirical range is 256–1024 tokens. Too small and chunks are semantically incomplete; too large and embeddings lose discriminative power while generation context gets wasted. Let evaluation data decide — don't guess.
- Parent-child chunking (LlamaIndex's Parent Document / Sentence Window) is an underrated value play: retrieve with small, precise chunks; generate with large, complete context — far cheaper to implement than semantic chunking.
- Anthropic's Contextual Retrieval (September 2024) deserves separate mention: use an LLM to generate 50–100 tokens of "locating context" per chunk ("This chunk is from ACME's 2023 Q2 earnings report…"), prepend it, then embed and index with BM25. Officially reported: retrieval failure rate down 49%, and down 67% when combined with reranking. The cost: one LLM call per chunk at indexing time (amortizable with prompt caching).
- Late Chunking is another angle: run the whole long document through a long-context embedding model first, then pool each chunk's vector at the chunking boundaries — every chunk's vector naturally carries full-document information. About a 3% average gain over naive chunking on long-document retrieval benchmarks — modest, but low-intrusion to implement.
4. Retrieval Techniques
Dense vector retrieval: semantic matching
Embed queries and documents into the same vector space and find nearest neighbors by cosine similarity/inner product. Strong at paraphrases, cross-language, and fuzzy intent; weak at exact matching — product SKUs, error codes, personal names, IDs, the "not one character can be wrong" queries are where vector retrieval routinely falls over.
BM25: the evergreen keyword matcher
BM25 is the classic TF-IDF-family retrieval algorithm, depending on no neural network at all. It's extremely strong at exact term matches and costs zero embedding compute. The 2026 reality: no serious production RAG should run vector-only retrieval.
Hybrid retrieval + RRF: the production standard
The industry-standard move: run BM25 and vector retrieval in parallel and fuse the two rankings with Reciprocal Rank Fusion (RRF):
python
def rrf_fuse(bm25_hits: list[str], vector_hits: list[str], k: int = 60) -> list[str]:
"""Reciprocal Rank Fusion: fuse by the reciprocal of each result's rank — no score-scale alignment needed."""
scores: dict[str, float] = {}
for rank, doc_id in enumerate(bm25_hits):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank + 1)
for rank, doc_id in enumerate(vector_hits):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank + 1)
return sorted(scores, key=scores.get, reverse=True)RRF's advantage is that it looks only at ranks, not scores — BM25 and vector-similarity scores live on completely different scales; a weighted sum needs tuning, RRF doesn't. Retrieval benchmarks in recent years repeatedly confirm: the two-stage "hybrid retrieval + neural rerank" pipeline stably beats any single-stage method.
Query rewriting and expansion
The user's raw question is often a poor retrieval query. Common techniques:
- Query rewriting: use an LLM to rewrite colloquial, pronoun-laden questions into self-contained retrieval queries. Near-mandatory in multi-turn conversations ("how much is it?" — the "it" must be resolved).
- Multi-query expansion: have the LLM generate 3–5 semantically equivalent but differently worded queries; retrieve in parallel and merge-dedupe. Trade a little latency for recall.
- HyDE (Hypothetical Document Embeddings, 2022): have the LLM "pretend to answer" first, generating a hypothetical answer, then use that answer (not the original question) for vector retrieval. The idea: answers sit closer to answers in vector space than questions do. Effective for zero-shot retrieval across wide domain gaps, but if the LLM's hypothetical answer is wildly wrong it drags retrieval off course — it's a recall booster, not a free lunch.
- Step-Back Prompting: abstract the specific question into a higher-level one ("which school did X attend in August 2023" → "X's education history") and retrieve both. Fits detail questions that need background knowledge.
Rewrite first, expand second, HyDE last
Query rewriting fixes "the query itself is ambiguous" — nearly every conversational RAG should do it; multi-query expansion fixes "insufficient recall" — cheap and reliably beneficial; HyDE fixes "the question and documents are semantically far apart" — scenario-dependent gains with a derailment risk. Add them one at a time in this order, validating each with an evaluation — don't switch them all on at once.
Query rewriting in multi-turn conversations is just one cheap LLM call (a small model suffices):
python
def contextualize_query(history: list[dict], question: str, llm) -> str:
"""Rewrite a context-dependent follow-up in a multi-turn conversation into a self-contained retrieval query."""
if not history: # no rewrite needed on the first turn
return question
prompt = (
"Given the conversation below and the user's latest question, rewrite the latest "
"question as a retrieval query that can be understood without the conversation history. "
"Output only the rewritten question.\n\n"
+ "\n".join(f"{m['role']}: {m['content']}" for m in history[-6:]) # only the last few turns
+ f"\n\nLatest question: {question}"
)
return llm.complete(prompt).strip()Note the "only the last few turns" — full history is expensive and imports off-topic interference. A small-parameter model is entirely sufficient for this step; don't spend a flagship model on it.
Reranking: the highest-ROI model call in the pipeline
Cross-encoder rerankers (Cohere Rerank, the BGE-Reranker family, etc.) take "query + candidate document" pairs as input and output relevance scores directly — far more precise than vector dot products. The price: no precomputation, only online pairwise scoring, so they're used solely to fine-rank the coarsely recalled top-k.
Rule of thumb: coarse vector ranking of top-50, rerank down to top-5 — the standard move that lifts retrieval from "usable" to "good." Recent benchmarks on mixed text+table documents also show two-stage combinations like "hybrid retrieval + Cohere Rerank" beating all single-stage methods by a clear margin.
5. Choosing a Vector Database
The 2026 consensus: there is no "best" vector database — only the one matching your scale and ops capacity. A rough decision compass first:
| Option | Type | Suited scale | Traits | Caveats |
|---|---|---|---|---|
| pgvector | Postgres extension | < ~10M vectors | Zero new infrastructure; vectors share transactions and SQL filtering/joins with business data | Latency and throughput lag dedicated engines at very large scale |
| Qdrant | Open source (Rust) + cloud | 10M–100M | Low latency, strong filtering, easy deployment | Smaller ecosystem than Milvus |
| Milvus | Open-source distributed + cloud | 100M+ | Designed for extreme scale; leads the four on GitHub stars (~44K) | Highest ops complexity; overkill at small scale |
| Weaviate | Open source + cloud | 10M–100M | Built-in hybrid search and modular vectorization; flexible schema | Heavier resource consumption |
| Pinecone | Fully managed SaaS | Unlimited | Zero ops, fastest to start | Monthly fees can reach tens of thousands of dollars at production scale; data-residency constraints |
| Chroma | Open-source embedded | Prototypes / small scale | Runs in-process; best for local experiments | Don't take it to production |
The pragmatic selection path is the one the industry keeps re-validating: already running Postgres and vector count is modest? Start with pgvector — one fewer component is one fewer failure class and one less ops burden. When latency or scale genuinely pushes the ceiling, migrate to Qdrant (self-hosted) or Pinecone/Weaviate Cloud (managed). Teams that start with a Milvus cluster are mostly paying an ops tax for a problem they don't have.
Know the index algorithm a little better than the brand
Mainstream vector databases are all HNSW (hierarchical navigable small world graphs) and variants underneath. You only need two knobs: ef (candidates explored at search time) tunes the recall/latency trade-off, and M (graph connectivity) tunes the memory/precision trade-off. Problems a database switch can't fix are often fixed by tuning these two.
6. Agentic RAG
Naive RAG is a fixed pipeline: retrieve once, generate once, and answer through gritted teeth no matter how good or bad the retrieval was. Agentic RAG turns retrieval into an tool of the agent, with the model deciding autonomously inside the Agent Loop:
┌─────────────────────────────────┐
│ Agent Loop │
│ │
User question ►│ Think: does this need retrieval?│
│ ├─ No → answer directly │
│ └─ Yes → call the search tool │
│ ↓ observe the result │
│ Is the result good enough? │
│ ├─ No → rewrite and retrieve again │
│ ├─ No → switch tools (SQL/API)│
│ └─ Yes → generate a cited answer │
└─────────────────────────────────┘The essential differences are three:
- Retrieval goes from "mandatory step" to "optional action." The model can judge "this question doesn't need the knowledge base," skipping pointless retrieval latency and noise.
- Multi-turn retrieval and self-correction. When results disappoint, the agent can rewrite the query, change keywords, switch data sources — instead of handing garbage to the generator in one shot. Grading retrieval relevance and deciding whether to retry is the classic structure of LangGraph's official Agentic RAG tutorial (see the LangGraph page).
- Retrieval is just one tool in the box. The agent can route and combine across the vector store, SQL, web search, and internal APIs — the right answer to multi-hop questions ("first look up A's department, then that department's budget").
The cost is equally real: every extra "think-retrieve" round is another LLM call. A January 2026 experimental paper systematically comparing RAG paradigms measured Agentic RAG costing up to 3.6× the naive approach, from extra reasoning steps and repeated tool calls. The conclusion is direct: simple factual QA → naive RAG; multi-hop, verification-needing, multi-source problems → Agentic RAG is worth it. This is also the canonical application of the "route by question complexity" strategy in Cost Optimization.
7. Evaluation
An unevaluated RAG system is running naked. Evaluation has two layers, both indispensable:
Retrieval-layer metrics: machine-computable
- Hit Rate@k: does the top-k contain the document holding the correct answer? The most intuitive quality gate.
- MRR (Mean Reciprocal Rank): the mean of the reciprocal rank of the first relevant result. A hit at position 10 versus a hit at position 1 affects downstream generation enormously — MRR captures exactly that.
- nDCG: a finer metric accounting for graded relevance; high annotation cost, and most teams are fine with the first two.
The prerequisite is a labeled test set: a set of "question → should-hit document" pairs. Even 50–100 entries will expose big problems. Without one, cold-start by having an LLM generate questions backward from your existing documents.
End-to-end metrics: LLM-as-a-judge
- Faithfulness: is every claim in the answer supported by the retrieved context? The direct measure of hallucination.
- Answer Relevancy: is the answer actually answering the user's question, rather than something adjacent?
- Context Precision / Recall: the share of retrieved context that's relevant / whether everything that should have been retrieved was.
RAGAS is the de facto standard framework at this layer: it uses an LLM judge to compute the metrics above automatically, needs no human-labeled answers, scores 0–1, and integrates directly with LangChain, LlamaIndex, and friends. The current RAGAS has evolved from a "RAG-specific evaluation library" into a general LLM-application evaluation framework with custom metrics and experiment management. Peers include DeepEval and TruLens; the differences are minor — getting something running matters more than which one you pick.
Evaluation scores are not a get-out-of-jail card
RAGAS's faithfulness comes from an LLM judge, and judges make mistakes — especially in specialized domains. The right posture: automated metrics for regression testing and trend monitoring (catching degradation), plus sampled human review before launch and major changes. Wire evaluation into CI and your observability platform so every prompt, chunk-parameter, and model change has a score to compare against — the general methodology is in Evaluation.
8. Common Failure Cases and a Troubleshooting Checklist
The classic failure points from Barnett et al.'s 2024 paper "Seven Failure Points When Engineering a RAG System" (based on three real case studies) remain the backbone of troubleshooting:
- Missing content: the answer isn't in the knowledge base at all, but the model answers anyway. → Add corpus + allow "I don't know" in the prompt.
- Missed top-ranked: the answer is in the corpus but didn't make top-k. → Tune chunk size/overlap, adopt hybrid retrieval, add query rewriting.
- Not in context: retrieved, but dropped from the final prompt by reranking or truncation. → Check the top-k → top-n pruning logic and the context budget.
- Not extracted: the answer is in the context but the model didn't use it — the classic symptom of an overlong, noisy context. → Remove irrelevant chunks, compress the context, place key evidence at the ends.
- Wrong format: asked for a table/JSON, got prose. → Strengthen instructions, add few-shot examples, add output validation with retry.
- Incorrect specificity: too vague or too rambling; granularity doesn't match the question. → Specify the desired level of detail in the prompt.
- Incomplete answers: the answer spans multiple chunks and only half got assembled. → Parent-child chunking, widen recall, agentic multi-turn retrieval.
When troubleshooting hands-on, walk this sequence before reaching for a new model:
- Isolate the fault domain first: take a failing case and inspect retrieval alone — is the problem retrieval or generation? High faithfulness but a wrong answer → retrieval's fault; the answer was in the context but the model missed it → generation's fault.
- Inspect parse artifacts: print the parsing and chunking output for the problem document and read it yourself. A large share of "retrieval is bad" is actually "parsing shredded the tables/multi-column layout."
- Iterate one variable at a time: change only one thing (chunk size, top-k, reranking on/off, rewrite strategy) and compare eval scores. Changing three at once changes nothing knowably.
- Look at distributions, not anecdotes: one failed case may be noise; after categorizing 20 failed cases, the repair priorities surface on their own.
One honest closing note: RAG's difficulty was never "standing it up" — any framework does it in fifty lines. The difficulty is the dirty work of grinding hit rate from 60% to 90%+: parsing, chunking, rewriting, evaluation. That's also why it's a high-frequency interview topic — see the Job-Search Knowledge Map for what those roles probe. To build one yourself, start from Build Your First Agent from Scratch.
References
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020) — the original RAG paper; the starting point for understanding the paradigm.
- Seven Failure Points When Engineering a RAG System (Barnett et al., 2024) — the failure-mode taxonomy from real case studies; the main basis for this page's troubleshooting checklist.
- Precise Zero-Shot Dense Retrieval without Relevance Labels (HyDE, 2022) — the original HyDE query-expansion paper.
- Introducing Contextual Retrieval (Anthropic, 2024) — the contextually enriched chunking scheme; reports retrieval failure rates down 49–67%.
- Ragas official documentation — metric definitions and current usage for the RAG evaluation framework.
- Traditional RAG vs. Agentic RAG (NVIDIA Developer Blog, 2025) — a comparison of the two paradigms and a selection perspective.
- Is Agentic RAG worth it? An experimental comparison (arXiv, 2026) — the experiment measuring Agentic RAG's cost at up to 3.6× the naive approach.
- Best Vector Databases in 2026: A Complete Comparison Guide (Firecrawl) — scale, ecosystem, and selection comparisons for the mainstream vector databases.