Skip to content

RAG: Retrieval-Augmented Generation

At a glance RARAG connects external knowledge to large models via "retrieve first, generate after," solving three major problems: knowledge cutoff, hallucination, and private data. This article breaks down RAG motivation, the index/retrieve/generate three-stage flow, Naive/Advanced/Modular evolution, retrievers and evaluation, and selection comparison with fine-tuning and long context.

This page contains time-sensitive content, current as of 2025-06; job descriptions, rankings, product features, and other information may have changed. Please verify with original sources before citing.

RAG: Retrieval-Augmented Generation ​

RAG (Retrieval-Augmented Generation) is an architectural pattern of "first retrieve external knowledge, then append the retrieved results to the context, and finally have the large model generate an answer." It turns the large model from "closed-book exam" to "open-book exam": the model doesn't need to memorize all facts, just learns to "answer based on given materials." RAG is the #1 choice for enterprise large model deployment, and the most effective, actionable solution for mitigating hallucination: causes and mitigation.

I. Motivation: Four Weaknesses of Large Models ​

WeaknessManifestationRAG's Solution
Knowledge cutoffDoesn't know new events/knowledge after trainingRetrieve from external knowledge base in real-time
HallucinationConfidently fabricates factsConstrains generation with retrieved real materials
Private dataCompany documents, databases not publicRetrieve from enterprise private knowledge base
Non-traceableAnswers have no basis, can't be verifiedAnswers link back to original snippets

RAG's underlying judgment is: "whether the model knows something" and "whether the model can answer based on materials" are two different things. The former relies on pretraining (expensive, static); the latter relies on context (cheap, updatable) — RAG chooses the latter.

The essence of RAG

RAG separates "knowledge" from model weights into an indexable, updatable, auditable external store. The cost: one extra retrieval step per answer. The gain: knowledge never expires, is always traceable, and private data never enters the model.

II. Three-Stage Flow: Index → Retrieve → Generate ​

┌────────────────────────────────────────────────────┐
│ ① Index stage (offline, one-time)                  │
│   Documents → parse/clean → chunking → embed → vector DB │
├────────────────────────────────────────────────────┤
│ ② Retrieve stage (online, per query)               │
│   Query → embed → similarity top-k → (rerank) → candidates │
├────────────────────────────────────────────────────┤
│ ③ Generate stage (online, per answer)              │
│   [system prompt + retrieved snippets + query] → LLM → cited answer │
└────────────────────────────────────────────────────┘

1. Index Stage ​

  • Parse and clean: PDF/Word/web → plain text, remove headers/footers, TOC, noise;
  • Chunking: cut long text into retrieval units, see below;
  • Embedding: use an embedding model (e.g., bge, E5, text-embedding series) to encode each chunk into a vector;
  • Storage: write to vector DB (FAISS, Milvus, pgvector, Qdrant, etc.), keeping original text for citation and reranking.

2. Retrieve Stage ​

Encode the user query into a vector, do similarity search against the DB (cosine/inner product), take top-k; high-quality systems do a reranking step to push the most relevant snippets to the front.

3. Generate Stage ​

Append retrieved snippets into the prompt:

System: You are a customer service assistant. Only answer based on "knowledge base materials." If the materials don't contain the answer, explicitly state so.
Materials:
[1] Return policy: within 7 days of receipt, no-reason returns are accepted...
[2] Refund timeline: 3–5 business days after return approval...
User: I returned something yesterday. When will I get my refund?
Answer: According to material [2], 3–5 business days after return approval...

4. Why Index Quality Sets the Ceiling ​

RAG has a "bucket effect": the system's ceiling is determined by its weakest link, and for most projects, that weak link is the index. Document parsing errors (tables split apart, OCR misreads on scans) pollute retrieval directly; poorly chunked text causes semantic breaks; embedding quality determines whether "semantically similar" can be recalled. Conversely, the ROI on the index stage is extremely high: a clean, reasonably chunked, metadata-rich index often yields more quality improvement than switching to a stronger generation model. This is why production RAG teams spend most of their time on document governance and index iteration rather than model tuning. Engineering details for the index/retrieve/generate stages are in RAG in Practice.

5. A Minimum Viable RAG Prototype ​

Don't be intimidated by the toolchain — a minimum prototype needs only four things: an embedding model (bge or OpenAI's embedding API), a vector DB (FAISS or an in-memory implementation suffices), some retrieval code, and a prompt-assembly snippet. Get "query → retrieve → assemble → generate → cite" working on 50 documents, then gradually add reranking, hybrid retrieval, and evaluation. Get it working first, optimize later — this is RAG project's #1 principle to avoid over-engineering.

III. Retrievers and Chunking: The Two Cornerstones of RAG Quality ​

1. Three Types of Retrievers ​

RetrieverPrincipleProsCons
Sparse BM25Lexical matching + TF-IDF weightingNo training needed, exact word match, interpretableNo semantics, poor synonym recall
Dense retrievalBi-encoder embedding cosine similaritySemantic matching, cross-languageDepends on embedding quality, needs training data
Hybrid retrievalBM25 + dense weighted fusion (RRF)Balances exact and semantic, most robustOne more engineering layer

Reranking typically uses a cross-encoder: concatenating the query with each candidate and scoring jointly, which is more precise than the bi-encoder's "encode separately then compare similarity" (but slower, so only applied to top-k for fine ranking).

When retrieval quality is poor, all downstream optimization (prompts, models) is "fine-tuning on wrong information." This is why professional teams make retrieval evaluation (Recall@k, MRR, nDCG) a standalone pipeline: get retrieval right first, then talk generation. A practical experience threshold: when top-5 recall is below 70%, optimize index and retrieval first, not generation. Evaluation tools and workflows are in Evaluations in Practice.

2. Chunk Strategy Comparison ​

StrategyApproachSuitable For
Fixed sizeCut at 256/512 tokens, with overlapGeneral, simple
Recursive characterCut by paragraph/sentence/word level fallbackComplex structured docs
Semantic chunkingUse embedding similarity to find natural boundariesLong docs, obvious topic shifts
Structure-awareCut by Markdown/HTML headings, table rowsManuals, web pages, tables

Rule of thumb: chunks that are too small lose context; too large introduce retrieval noise — generally experiment in the 200–800 token range. More engineering detail in RAG in Practice.

IV. Evolution Path: Naive → Advanced → Modular ​

1. Naive RAG (2020–early 2023) ​

The original paper, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020, Facebook AI Research), used DPR retrieval + BART generation, achieving SOTA at the time on open-domain QA (Natural Questions, etc.). Early practice's "naive version" was the most direct "vector retrieval + assemble + generate." Typical problems: poor retrieval quality, noisy snippets polluting answers, no reranking, no query understanding — if retrieval is wrong, even the strongest generation won't save it.

2. Advanced RAG (2023) ​

Patching three stages separately:

StageOptimization
Query sideQuery rewriting, multi-query expansion, HyDE (generate hypothetical answer first, then retrieve)
Index sideBetter chunking, metadata filtering, parent-child chunking
Retrieve sideHybrid retrieval + reranking + context compression (keep only relevant snippets)

Representative: HyDE (2022, uses "generative pseudo-documents" to improve zero-shot retrieval), RAG-Fusion (multi-query RRF).

3. Modular RAG (late 2023–2024) ​

Breaking RAG into orchestratable modules (query planning, routing, memory, reflection, rewriting), allowing retrieval to loop multiple times, and even letting the model self-evaluate retrieval results:

  • Self-RAG (2023): the model generates "reflection tokens," self-judging "whether retrieval is needed, whether the result is relevant, whether the answer is supported by evidence";
  • CRAG (Corrective RAG) (2024): scores retrieval result quality, triggers "query rewrite and re-retrieve" if poor, avoiding answering with flawed materials.

Modular RAG is now the de facto form of enterprise RAG: RAG is no longer a fixed pipeline but a programmable retrieval-generation loop.

The main thread of RAG evolution is actually an increase in "cognitive complexity": Naive stage is "fixed pipeline" (one retrieval, one generation), Advanced stage is "controllable pipeline" (each stage can be optimized), and Modular stage is "programmable pipeline" (modules freely orchestrated). This thread is highly isomorphic with Agent development — in fact, Modular RAG is already deeply integrated with LLM-Based Agents: retrieval, reranking, rewriting, reflection can all be Agent tools and decisions. The boundary between RAG and Agent will blur further; understanding this trend helps choose technical direction.

Three Naive RAG failure modes

① Retrieving irrelevant snippets → answer misdirected; ② Too many snippets flood context → key info drowned; ③ Knowledge base updated but vector DB not rebuilt → old data answering. If "RAG performs poorly," check these three first before switching models.

V. RAG vs Fine-tuning vs Long Context ​

DimensionRAGFine-tuningLong Context
Knowledge sourceExternal retrieval (can update in real-time)Written into weights (static)Full text in context
Knowledge cutoffUpdate the indexRequires retrainingMust put all new docs in context
Hallucination mitigationStrong (evidence-based)Weak (still fabricates)Medium (long text still "lost in the middle")
TraceabilityCan cite originalNoNo
CostRetrieval overhead + context per requestOne-time training costHigh long-text inference cost (O(n²) attention)
LatencyMedium (one extra retrieval step)LowHigh (long input encoding)
Suitable forFactual, knowledge base Q&A, real-time dataStyle/format/domain behaviorSingle-document deep analysis, codebase context
Combined use✓ Most common combo: RAG gives knowledge + fine-tuning gives style

Selection mantra: knowledge problems → RAG; behavior problems → fine-tuning; single very-long-document deep analysis → long context. The three aren't mutually exclusive but additive — see Fine-tuning: SFT and Parameter-Efficient Fine-tuning and Context and Long Context.

There's a more nuanced view of their combination: they solve problems at different levels. RAG solves "where does knowledge come from" (fact source), fine-tuning solves "what does behavior look like" (style and format), and long context solves "how much can be seen at once" (information bandwidth). A common optimal solution: use RAG to access the latest and private knowledge, use long context to carry single large documents (e.g., contracts, papers), and use fine-tuning or prompts to solidify output style. The recommended decision order: RAG first, then long context (as needed), then fine-tuning (when there's a clear behavior requirement) — this order also follows the engineering principle of "cheap first, expensive later; reversible first, irreversible later."

One comparison dimension worth emphasizing separately: update cost. One fine-tuning costs tens of thousands of yuan and takes days; long context sends the full text to the model every time (high token cost); while RAG updates only require rebuilding the index (minutes to hours, low cost). In scenarios where "knowledge changes frequently" (policies, pricing, inventory), RAG is virtually the only realistic choice; in "knowledge is stable and format requirements are high" scenarios, fine-tuning may be more suitable. Adding the "update frequency" dimension to the decision is why many teams choose wrong.

Why hasn't long context "killed" RAG?

Long context solves "can it read long," but not "where does knowledge come from": enterprise docs update daily, you can't stuff every one into a 128K context, and longer means more expensive and more prone to "lost in the middle." RAG solves both cost and precision by "bringing only relevant snippets."

VI. RAG Evaluation: Three Layers — Retrieval, Generation, End-to-End ​

LayerMetricsNotes
RetrievalRecall@k, MRR, nDCGAre relevant snippets recalled? Is ranking high?
GenerationFaithfulness, answer relevanceIs the answer strictly based on retrieval snippets? On-topic?
End-to-endRAGAS composite score, human eval, A/BOverall quality perceivable by users

Three common RAG evaluation misconceptions. First, only looking at end-to-end scores — a high end-to-end score doesn't mean retrieval is good; the generation model may have "guessed right," masking retrieval flaws — also look at retrieval-layer metrics (Recall@k, MRR). Second, evaluation set misaligned with production distribution — measuring your knowledge base on public QA sets from papers is basically meaningless; the evaluation set must be sampled from real user queries. Third, ignoring "negative samples" — only testing "answers in the DB" overestimates the system; test "questions not in the DB" and see whether the model honestly says "I don't know" (rather than fabricating). These three correspond to the three principles of RAG evaluation: "layered, real, comprehensive coverage." Detailed methods in Evaluations in Practice and Evaluation and Benchmarks.

RAGAS is a framework evaluation proposed in 2023 (paper in references), breaking RAG quality into "faithfulness + answer relevance + context relevance" across three metrics, enabling low-cost automated scoring. The evaluation set must be custom-built: sample from real user queries, annotate "expected snippets + expected answers." See Evaluations in Practice.

VII. Representative Systems ​

SystemTimeHighlights
Microsoft Bing Chat (New Bing)2023.2First large-scale integration of search engine + generative answer, Q&A with source links
Google Bard / Gemini2023.3 onwardSearch-augmented generation, evolving progressively
Perplexity AI2023 onwardPure RAG-style AI search engine, answers with full source citations
Bespoke-Minerva2023.12Arcee open-sourced "RAG-specialized fine-tuned model" (based on Llama), demonstrating "retrieval + fine-tuning" combination
Enterprise platforms2023 onwardAzure AI Search, AWS Kendra, OpenSearch, etc. productizing RAG

Representative systems validate one conclusion: RAG's moat isn't the model, it's "knowledge base + retrieval quality + citation experience" — whichever team does best on document governance, index quality, and retrieval tuning has the most reliable Q&A product.

Search engines are RAG's most natural host — they already have large-scale indexing, ranking, and click feedback. Bing Chat and Google Bard combine "search results + generative answers," improving both information consumption experience and providing infrastructure for "citation traceability." This explains why RAG benchmark systems mostly come from search vendors: RAG's difficulty has never been on the generation side, but on the engineering accumulation of the retrieval side.

VIII. RAG Engineering Selection ​

1. Vector Database Selection ​

OptionDeploymentScaleSuitable Scenario
FAISSEmbedded in applicationTens of millionsPrototypes, single-machine, extreme speed
MilvusStandalone service (distributable)100M+Production, high concurrency
pgvectorPostgreSQL pluginMillionsSame stack as business DB, transactional consistency
QdrantStandalone serviceTens of millionsFull-featured, easy to use
ElasticsearchStandalone service100M+Existing ES ecosystem, hybrid retrieval

Selection decision points: concurrency, data volume, ops cost, need for hybrid retrieval. Engineering landing in RAG in Practice and Framework and Tool Selection.

2. Embedding Model Selection ​

ModelDimensionsCharacteristics
bge (BAAI)1024Strong Chinese, multilingual
E5 (Microsoft)1024Multilingual, high retrieval benchmark
text-embedding-3 (OpenAI)1536/3072Closed-source API, good ecosystem
GTE (Alibaba)1024Strong Chinese

Watch out when selecting: dimensions affect storage cost and retrieval speed; context length determines the single-chunk encoding limit; models for different languages shouldn't be mixed.

3. A Complete RAG Prompt Template ​

System prompt:
You are an assistant that strictly answers based on "the knowledge base." Rules:
1. Only use information from the materials; do not fabricate;
2. Mark the source number after each claim, e.g. [1];
3. If the materials are insufficient to answer, explicitly say "no relevant information in the knowledge base," and provide the closest retrieval suggestion.

Materials:
[1] {retrieved snippet 1}
[2] {retrieved snippet 2}
[3] {retrieved snippet 3}

User question: {query}

4. Chunk Parameter Experimentation Method ​

ParameterStarting PointAdjustment Direction
Chunk size400 tokensAnswer snippets are long → increase; too much noise → decrease
Overlap50–100 tokensKey sentences cut → increase
Top-k4Not enough recall → increase; too much noise → decrease
RerankingMust enablePoor relevance → switch cross-encoder

5. Retrieval Tuning Checklist ​

  • [ ] Test set covers real query distribution (golden set)
  • [ ] Hybrid retrieval (BM25 + vector) is enabled
  • [ ] Reranking model is connected and effect is compared pre/post
  • [ ] Query rewriting / multi-query experimented
  • [ ] Access control (ACL) takes effect before retrieval
  • [ ] Evaluation metrics (Recall@k, faithfulness, answer relevance) are live

6. Organization-Level RAG Landing Advice ​

The gap between enterprise RAG and personal prototype RAG is almost entirely in "governance" not "algorithm": permission models (who can see which documents) must be enforced before retrieval, otherwise it's a data leak incident; knowledge updates need clear ownership and approval flows; vector DB versions must link with document versions; answers need traceable citations and audit paths. Another undervalued aspect is cold-start evaluation: with zero online data, first build a golden set of 100 real human-organized questions, then automate gradually. These suggestions and the more complete landing framework are in RAG in Practice and Common Pitfalls and Anti-Patterns.

Beyond the checklist, there's a "soft" factor that determines success or failure: document governance. The same entity with multiple spellings ("OpenAI" and "openai," full company name vs. abbreviation), contradictory statements from different sources, frequently changing versions — all of these cause retrieval and generation to repeatedly fail. Governance investment doesn't directly show up in code, but it's the ceiling of retrieval quality. The relationship between document governance and RAG is like data cleaning to pretraining — methodology in Pretraining: Data and Objectives.

The #1 principle of RAG tuning

Fix retrieval first, then generation. 80% of RAG quality problems come from "the retrieved stuff is wrong," not "the model can't answer." Retrieval tuning methods (HyDE, reranking, hybrid retrieval) in RAG in Practice.

IX. Common Failure Mode Reference ​

Failure ModeSymptomRoot CauseCountermeasure
Retrieval off-targetAnswer unrelated to queryQuery and doc semantics misalignedQuery rewriting, HyDE, hybrid retrieval
Context drowningRelevant snippets drowned by noiseChunk too large / top-k too highSmaller chunks, reranking, compression
Low faithfulnessAnswer fabricates detailsModel free-playStrong constraint prompts, citation format, evaluation fallback
Stale knowledgeOld answersIndex not updatedIncremental indexing, freshness filtering
Private data leakAnswer leaks unrelated docsMissing permission filterACL filtering before retrieval

A final piece of experience: RAG project failures, nine out of ten times the problem is data and retrieval, not the model — first get document governance and retrieval quality to 90 points, and the generation side naturally stabilizes.

X. RAG Advanced Topics and FAQ ​

1. Multi-hop Retrieval: What If the Answer Is in Two Places? ​

Single-turn retrieval can only get snippets "directly similar to the query," while many real problems need a reasoning chain:

Query: Where is the headquarters of that chip company acquired last week?
Hop 1: Retrieve "last week chip company acquisition" → find "XX company acquired by YY"
Hop 2: From XX company name, retrieve "XX company headquarters" → find the answer

Multi-hop retrieval implementation: iterative retrieval (using previous answer as next query), query expansion (retrieve multiple sub-questions simultaneously), knowledge graph enhancement (using entity relation graphs to bridge cross-document associations). Higher complexity, but covers "indirect association" problems.

2. Caching and Incremental Updates ​

StrategyApproachBenefit
Vector cacheSame/similar queries directly hit historical answersSave 80%+ call cost
Incremental indexingDocument changes only rebuild changed chunksSecond-level updates
Dual-zone indexingHot zone (frequently asked) + cold zone (full)Control cost
Freshness filteringFilter by time/version at retrievalAvoid old data answering

How "alive" the knowledge base is often impacts RAG quality more than model selection.

3. Agentic RAG: Handing Retrieval to Agents ​

Traditional RAG is "one retrieval + one generation"; Agentic RAG lets LLM-Based Agents autonomously decide: which DB to query, how many times, is the result sufficient, should the strategy change. Typical forms:

  • Routing: first judge question type, then choose the corresponding knowledge base;
  • Tool-based: make retrieval, reranking, and web search tools, orchestrated by the Agent;
  • Reflection: evaluate generation using self-evaluation, re-retrieve if insufficient (Self-RAG approach).

Agentic RAG is flexible but more expensive and less stable — suitable for "high-complexity problems needing multi-source info" scenarios.

4. Automating RAG Evaluation ​

Offline pipeline (run every change):
① Build golden set: 50–200 real questions + expected answers + expected citations
② Run retrieval: calculate Recall@k, MRR
③ Run generation: score faithfulness/relevance via RAGAS or LLM judge
④ Compare baseline (previous version) → regression report
⑤ Small-traffic online A/B validation

Metric and tool details in Evaluations in Practice.

5. FAQ Quick Answers ​

QuestionQuick Answer
Can RAG eliminate hallucination?Not completely, but can significantly reduce factual hallucination
Fine-tuning or RAG?Knowledge → RAG, behavior → fine-tuning, can be combined
Can long context replace RAG?No, cost and update speed don't allow it
How large a vector DB is needed?100K-level: pgvector/FAISS sufficient
Which embedding for Chinese?bge / GTE / E5 Chinese versions
Where does RAG most commonly fail?Retrieval quality, then chunk strategy

Three-stage judgment for RAG

Naive can run, Advanced can improve scores, Modular/Agentic handles complex business. Don't jump to the heaviest architecture first — get index quality and evaluation sets right, and most projects already win half the battle.

Finally, viewing RAG in a larger context: it and "context engineering" are two sides of the same coin. Whether RAG, long context, or prompt compression, the essence is delivering "the most relevant information" to the model within limited inference budget. The only difference is "who does the filtering": RAG uses a retriever, long context lets the model search itself in large input, prompt compression uses a summarizer. Understanding this unified perspective lets you flexibly combine in specific scenarios: short docs → full text directly; long knowledge base → RAG; ultra-long single doc → RAG + long context hybrid. Tech buzzwords expire, but the proposition of "efficiently delivering relevant information to the model" does not.

One more note: the "division of labor" between retrieval and generation is fundamentally a choice between "externalized knowledge vs. internalized knowledge" — as long as knowledge changes, RAG deserves to remain the default option.

XI. RAG Variants and Adjacent Technologies ​

1. Main Variant Comparison ​

VariantCore IdeaSuitable ForCost
Standard RAGVector retrieval + generationGeneral knowledge baseRetrieval quality sets ceiling
GraphRAGKnowledge graph + graph traversal retrievalStrongly associated, multi-hop problemsHigh graph construction cost
RAPTORRecursive clustering summary treeHierarchical understanding of long docsComplex offline construction
HyDEGenerate hypothetical answer first, then retrieveZero-shot retrieval cold startOne extra generation step
RAG-FusionMulti-query + RRF fusionImprove recallQuery cost doubles
Self-RAGSelf-reflection on whether to retrieveReduce unnecessary retrievalNeeds training reflection tokens

2. RAG and Knowledge Graph QA ​

Knowledge graph QA (KGQA) answers "who, where, related to whom" problems with structured triples — high certainty but high construction and maintenance cost; RAG is flexible but lower certainty. In practice, "graph + RAG" hybrid is common: use the graph to bridge associations, use vector retrieval for semantics.

3. RAG + Fine-tuning Combination Patterns ​

CombinationSolves WhatHow to Stack
RAG gives knowledge + fine-tuning gives styleKnowledge + customer service tone/formatRetrieval snippets + LoRA model
Fine-tune retrieverPoor retrieval qualityFine-tune bi-encoder/reranker with annotated pairs
Fine-tune generator for citation formatMessy citation formatFine-tune with citation-format instruction data

Combination principle: fine-tuning doesn't change knowledge (that's RAG's job), RAG doesn't change style (that's fine-tuning's job).

4. RAG and MCP / Toolification ​

RAG retrieval itself can be made into a tool: expose "retrieve enterprise knowledge base" via MCP, letting LLM-Based Agents call it on demand during tasks — the standard form of Agentic RAG.

A prompt for researchers: RAG's open problems concentrate in three areas — "when to retrieve" (query intent judgment), "when is retrieval enough" (relevance measurement and stopping conditions), and "how to pass retrieval result credibility to the generation side" (citation and uncertainty propagation). These three correspond to Modular RAG's routing, reflection, and generation control modules, and are the core of Self-RAG, CRAG, and related work. Understanding these three open questions grasps RAG's research thread more than chasing new frameworks.

XII. RAG Key Paper Checklist ​

Paper / WorkYearOne-Sentence
Retrieval-Augmented Generation (RAG)2020Founding work
DPR2020Dense retrieval bi-encoder
HyDE2022Zero-shot dense retrieval improvement
Self-RAG2023Self-reflective retrieval
CRAG2024Corrective retrieval quality
RAGAS2023Automated evaluation framework
GraphRAG2024Graph-augmented retrieval

For newcomers, the recommended paper reading order: first read the 2020 RAG original paper to build the overall "retrieval-generation" picture, then read DPR to understand dense retrieval, then look at HyDE, Self-RAG, CRAG for evolution logic, and finally use RAGAS to understand evaluation. For each, just focus on "what problem was solved, what mechanism was introduced" — no need to derive every line. Paper deep-dive methods in Core Paper Deep Dives.

One sentence for RAG tech selection

The simpler the need, the more basic the solution: 80% of Q&A can use "standard RAG + reranking + evaluation loop"; GraphRAG and Agentic RAG are only worth introducing when multi-hop or autonomous decision-making is explicitly needed.

XIII. Further Reading ​

References ​