Theme
RAG: Retrieval-Augmented Generation
RAG (Retrieval-Augmented Generation) is an architectural pattern of "first retrieve external knowledge, then append the retrieved results to the context, and finally have the large model generate an answer." It turns the large model from "closed-book exam" to "open-book exam": the model doesn't need to memorize all facts, just learns to "answer based on given materials." RAG is the #1 choice for enterprise large model deployment, and the most effective, actionable solution for mitigating hallucination: causes and mitigation.
I. Motivation: Four Weaknesses of Large Models
| Weakness | Manifestation | RAG's Solution |
|---|---|---|
| Knowledge cutoff | Doesn't know new events/knowledge after training | Retrieve from external knowledge base in real-time |
| Hallucination | Confidently fabricates facts | Constrains generation with retrieved real materials |
| Private data | Company documents, databases not public | Retrieve from enterprise private knowledge base |
| Non-traceable | Answers have no basis, can't be verified | Answers link back to original snippets |
RAG's underlying judgment is: "whether the model knows something" and "whether the model can answer based on materials" are two different things. The former relies on pretraining (expensive, static); the latter relies on context (cheap, updatable) — RAG chooses the latter.
The essence of RAG
RAG separates "knowledge" from model weights into an indexable, updatable, auditable external store. The cost: one extra retrieval step per answer. The gain: knowledge never expires, is always traceable, and private data never enters the model.
II. Three-Stage Flow: Index → Retrieve → Generate
┌────────────────────────────────────────────────────┐
│ ① Index stage (offline, one-time) │
│ Documents → parse/clean → chunking → embed → vector DB │
├────────────────────────────────────────────────────┤
│ ② Retrieve stage (online, per query) │
│ Query → embed → similarity top-k → (rerank) → candidates │
├────────────────────────────────────────────────────┤
│ ③ Generate stage (online, per answer) │
│ [system prompt + retrieved snippets + query] → LLM → cited answer │
└────────────────────────────────────────────────────┘1. Index Stage
- Parse and clean: PDF/Word/web → plain text, remove headers/footers, TOC, noise;
- Chunking: cut long text into retrieval units, see below;
- Embedding: use an embedding model (e.g., bge, E5, text-embedding series) to encode each chunk into a vector;
- Storage: write to vector DB (FAISS, Milvus, pgvector, Qdrant, etc.), keeping original text for citation and reranking.
2. Retrieve Stage
Encode the user query into a vector, do similarity search against the DB (cosine/inner product), take top-k; high-quality systems do a reranking step to push the most relevant snippets to the front.
3. Generate Stage
Append retrieved snippets into the prompt:
System: You are a customer service assistant. Only answer based on "knowledge base materials." If the materials don't contain the answer, explicitly state so.
Materials:
[1] Return policy: within 7 days of receipt, no-reason returns are accepted...
[2] Refund timeline: 3–5 business days after return approval...
User: I returned something yesterday. When will I get my refund?
Answer: According to material [2], 3–5 business days after return approval...4. Why Index Quality Sets the Ceiling
RAG has a "bucket effect": the system's ceiling is determined by its weakest link, and for most projects, that weak link is the index. Document parsing errors (tables split apart, OCR misreads on scans) pollute retrieval directly; poorly chunked text causes semantic breaks; embedding quality determines whether "semantically similar" can be recalled. Conversely, the ROI on the index stage is extremely high: a clean, reasonably chunked, metadata-rich index often yields more quality improvement than switching to a stronger generation model. This is why production RAG teams spend most of their time on document governance and index iteration rather than model tuning. Engineering details for the index/retrieve/generate stages are in RAG in Practice.
5. A Minimum Viable RAG Prototype
Don't be intimidated by the toolchain — a minimum prototype needs only four things: an embedding model (bge or OpenAI's embedding API), a vector DB (FAISS or an in-memory implementation suffices), some retrieval code, and a prompt-assembly snippet. Get "query → retrieve → assemble → generate → cite" working on 50 documents, then gradually add reranking, hybrid retrieval, and evaluation. Get it working first, optimize later — this is RAG project's #1 principle to avoid over-engineering.
III. Retrievers and Chunking: The Two Cornerstones of RAG Quality
1. Three Types of Retrievers
| Retriever | Principle | Pros | Cons |
|---|---|---|---|
| Sparse BM25 | Lexical matching + TF-IDF weighting | No training needed, exact word match, interpretable | No semantics, poor synonym recall |
| Dense retrieval | Bi-encoder embedding cosine similarity | Semantic matching, cross-language | Depends on embedding quality, needs training data |
| Hybrid retrieval | BM25 + dense weighted fusion (RRF) | Balances exact and semantic, most robust | One more engineering layer |
Reranking typically uses a cross-encoder: concatenating the query with each candidate and scoring jointly, which is more precise than the bi-encoder's "encode separately then compare similarity" (but slower, so only applied to top-k for fine ranking).
When retrieval quality is poor, all downstream optimization (prompts, models) is "fine-tuning on wrong information." This is why professional teams make retrieval evaluation (Recall@k, MRR, nDCG) a standalone pipeline: get retrieval right first, then talk generation. A practical experience threshold: when top-5 recall is below 70%, optimize index and retrieval first, not generation. Evaluation tools and workflows are in Evaluations in Practice.
2. Chunk Strategy Comparison
| Strategy | Approach | Suitable For |
|---|---|---|
| Fixed size | Cut at 256/512 tokens, with overlap | General, simple |
| Recursive character | Cut by paragraph/sentence/word level fallback | Complex structured docs |
| Semantic chunking | Use embedding similarity to find natural boundaries | Long docs, obvious topic shifts |
| Structure-aware | Cut by Markdown/HTML headings, table rows | Manuals, web pages, tables |
Rule of thumb: chunks that are too small lose context; too large introduce retrieval noise — generally experiment in the 200–800 token range. More engineering detail in RAG in Practice.
IV. Evolution Path: Naive → Advanced → Modular
1. Naive RAG (2020–early 2023)
The original paper, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020, Facebook AI Research), used DPR retrieval + BART generation, achieving SOTA at the time on open-domain QA (Natural Questions, etc.). Early practice's "naive version" was the most direct "vector retrieval + assemble + generate." Typical problems: poor retrieval quality, noisy snippets polluting answers, no reranking, no query understanding — if retrieval is wrong, even the strongest generation won't save it.
2. Advanced RAG (2023)
Patching three stages separately:
| Stage | Optimization |
|---|---|
| Query side | Query rewriting, multi-query expansion, HyDE (generate hypothetical answer first, then retrieve) |
| Index side | Better chunking, metadata filtering, parent-child chunking |
| Retrieve side | Hybrid retrieval + reranking + context compression (keep only relevant snippets) |
Representative: HyDE (2022, uses "generative pseudo-documents" to improve zero-shot retrieval), RAG-Fusion (multi-query RRF).
3. Modular RAG (late 2023–2024)
Breaking RAG into orchestratable modules (query planning, routing, memory, reflection, rewriting), allowing retrieval to loop multiple times, and even letting the model self-evaluate retrieval results:
- Self-RAG (2023): the model generates "reflection tokens," self-judging "whether retrieval is needed, whether the result is relevant, whether the answer is supported by evidence";
- CRAG (Corrective RAG) (2024): scores retrieval result quality, triggers "query rewrite and re-retrieve" if poor, avoiding answering with flawed materials.
Modular RAG is now the de facto form of enterprise RAG: RAG is no longer a fixed pipeline but a programmable retrieval-generation loop.
The main thread of RAG evolution is actually an increase in "cognitive complexity": Naive stage is "fixed pipeline" (one retrieval, one generation), Advanced stage is "controllable pipeline" (each stage can be optimized), and Modular stage is "programmable pipeline" (modules freely orchestrated). This thread is highly isomorphic with Agent development — in fact, Modular RAG is already deeply integrated with LLM-Based Agents: retrieval, reranking, rewriting, reflection can all be Agent tools and decisions. The boundary between RAG and Agent will blur further; understanding this trend helps choose technical direction.
Three Naive RAG failure modes
① Retrieving irrelevant snippets → answer misdirected; ② Too many snippets flood context → key info drowned; ③ Knowledge base updated but vector DB not rebuilt → old data answering. If "RAG performs poorly," check these three first before switching models.
V. RAG vs Fine-tuning vs Long Context
| Dimension | RAG | Fine-tuning | Long Context |
|---|---|---|---|
| Knowledge source | External retrieval (can update in real-time) | Written into weights (static) | Full text in context |
| Knowledge cutoff | Update the index | Requires retraining | Must put all new docs in context |
| Hallucination mitigation | Strong (evidence-based) | Weak (still fabricates) | Medium (long text still "lost in the middle") |
| Traceability | Can cite original | No | No |
| Cost | Retrieval overhead + context per request | One-time training cost | High long-text inference cost (O(n²) attention) |
| Latency | Medium (one extra retrieval step) | Low | High (long input encoding) |
| Suitable for | Factual, knowledge base Q&A, real-time data | Style/format/domain behavior | Single-document deep analysis, codebase context |
| Combined use | ✓ Most common combo: RAG gives knowledge + fine-tuning gives style |
Selection mantra: knowledge problems → RAG; behavior problems → fine-tuning; single very-long-document deep analysis → long context. The three aren't mutually exclusive but additive — see Fine-tuning: SFT and Parameter-Efficient Fine-tuning and Context and Long Context.
There's a more nuanced view of their combination: they solve problems at different levels. RAG solves "where does knowledge come from" (fact source), fine-tuning solves "what does behavior look like" (style and format), and long context solves "how much can be seen at once" (information bandwidth). A common optimal solution: use RAG to access the latest and private knowledge, use long context to carry single large documents (e.g., contracts, papers), and use fine-tuning or prompts to solidify output style. The recommended decision order: RAG first, then long context (as needed), then fine-tuning (when there's a clear behavior requirement) — this order also follows the engineering principle of "cheap first, expensive later; reversible first, irreversible later."
One comparison dimension worth emphasizing separately: update cost. One fine-tuning costs tens of thousands of yuan and takes days; long context sends the full text to the model every time (high token cost); while RAG updates only require rebuilding the index (minutes to hours, low cost). In scenarios where "knowledge changes frequently" (policies, pricing, inventory), RAG is virtually the only realistic choice; in "knowledge is stable and format requirements are high" scenarios, fine-tuning may be more suitable. Adding the "update frequency" dimension to the decision is why many teams choose wrong.
Why hasn't long context "killed" RAG?
Long context solves "can it read long," but not "where does knowledge come from": enterprise docs update daily, you can't stuff every one into a 128K context, and longer means more expensive and more prone to "lost in the middle." RAG solves both cost and precision by "bringing only relevant snippets."
VI. RAG Evaluation: Three Layers — Retrieval, Generation, End-to-End
| Layer | Metrics | Notes |
|---|---|---|
| Retrieval | Recall@k, MRR, nDCG | Are relevant snippets recalled? Is ranking high? |
| Generation | Faithfulness, answer relevance | Is the answer strictly based on retrieval snippets? On-topic? |
| End-to-end | RAGAS composite score, human eval, A/B | Overall quality perceivable by users |
Three common RAG evaluation misconceptions. First, only looking at end-to-end scores — a high end-to-end score doesn't mean retrieval is good; the generation model may have "guessed right," masking retrieval flaws — also look at retrieval-layer metrics (Recall@k, MRR). Second, evaluation set misaligned with production distribution — measuring your knowledge base on public QA sets from papers is basically meaningless; the evaluation set must be sampled from real user queries. Third, ignoring "negative samples" — only testing "answers in the DB" overestimates the system; test "questions not in the DB" and see whether the model honestly says "I don't know" (rather than fabricating). These three correspond to the three principles of RAG evaluation: "layered, real, comprehensive coverage." Detailed methods in Evaluations in Practice and Evaluation and Benchmarks.
RAGAS is a framework evaluation proposed in 2023 (paper in references), breaking RAG quality into "faithfulness + answer relevance + context relevance" across three metrics, enabling low-cost automated scoring. The evaluation set must be custom-built: sample from real user queries, annotate "expected snippets + expected answers." See Evaluations in Practice.
VII. Representative Systems
| System | Time | Highlights |
|---|---|---|
| Microsoft Bing Chat (New Bing) | 2023.2 | First large-scale integration of search engine + generative answer, Q&A with source links |
| Google Bard / Gemini | 2023.3 onward | Search-augmented generation, evolving progressively |
| Perplexity AI | 2023 onward | Pure RAG-style AI search engine, answers with full source citations |
| Bespoke-Minerva | 2023.12 | Arcee open-sourced "RAG-specialized fine-tuned model" (based on Llama), demonstrating "retrieval + fine-tuning" combination |
| Enterprise platforms | 2023 onward | Azure AI Search, AWS Kendra, OpenSearch, etc. productizing RAG |
Representative systems validate one conclusion: RAG's moat isn't the model, it's "knowledge base + retrieval quality + citation experience" — whichever team does best on document governance, index quality, and retrieval tuning has the most reliable Q&A product.
Search engines are RAG's most natural host — they already have large-scale indexing, ranking, and click feedback. Bing Chat and Google Bard combine "search results + generative answers," improving both information consumption experience and providing infrastructure for "citation traceability." This explains why RAG benchmark systems mostly come from search vendors: RAG's difficulty has never been on the generation side, but on the engineering accumulation of the retrieval side.
VIII. RAG Engineering Selection
1. Vector Database Selection
| Option | Deployment | Scale | Suitable Scenario |
|---|---|---|---|
| FAISS | Embedded in application | Tens of millions | Prototypes, single-machine, extreme speed |
| Milvus | Standalone service (distributable) | 100M+ | Production, high concurrency |
| pgvector | PostgreSQL plugin | Millions | Same stack as business DB, transactional consistency |
| Qdrant | Standalone service | Tens of millions | Full-featured, easy to use |
| Elasticsearch | Standalone service | 100M+ | Existing ES ecosystem, hybrid retrieval |
Selection decision points: concurrency, data volume, ops cost, need for hybrid retrieval. Engineering landing in RAG in Practice and Framework and Tool Selection.
2. Embedding Model Selection
| Model | Dimensions | Characteristics |
|---|---|---|
| bge (BAAI) | 1024 | Strong Chinese, multilingual |
| E5 (Microsoft) | 1024 | Multilingual, high retrieval benchmark |
| text-embedding-3 (OpenAI) | 1536/3072 | Closed-source API, good ecosystem |
| GTE (Alibaba) | 1024 | Strong Chinese |
Watch out when selecting: dimensions affect storage cost and retrieval speed; context length determines the single-chunk encoding limit; models for different languages shouldn't be mixed.
3. A Complete RAG Prompt Template
System prompt:
You are an assistant that strictly answers based on "the knowledge base." Rules:
1. Only use information from the materials; do not fabricate;
2. Mark the source number after each claim, e.g. [1];
3. If the materials are insufficient to answer, explicitly say "no relevant information in the knowledge base," and provide the closest retrieval suggestion.
Materials:
[1] {retrieved snippet 1}
[2] {retrieved snippet 2}
[3] {retrieved snippet 3}
User question: {query}4. Chunk Parameter Experimentation Method
| Parameter | Starting Point | Adjustment Direction |
|---|---|---|
| Chunk size | 400 tokens | Answer snippets are long → increase; too much noise → decrease |
| Overlap | 50–100 tokens | Key sentences cut → increase |
| Top-k | 4 | Not enough recall → increase; too much noise → decrease |
| Reranking | Must enable | Poor relevance → switch cross-encoder |
5. Retrieval Tuning Checklist
- [ ] Test set covers real query distribution (golden set)
- [ ] Hybrid retrieval (BM25 + vector) is enabled
- [ ] Reranking model is connected and effect is compared pre/post
- [ ] Query rewriting / multi-query experimented
- [ ] Access control (ACL) takes effect before retrieval
- [ ] Evaluation metrics (Recall@k, faithfulness, answer relevance) are live
6. Organization-Level RAG Landing Advice
The gap between enterprise RAG and personal prototype RAG is almost entirely in "governance" not "algorithm": permission models (who can see which documents) must be enforced before retrieval, otherwise it's a data leak incident; knowledge updates need clear ownership and approval flows; vector DB versions must link with document versions; answers need traceable citations and audit paths. Another undervalued aspect is cold-start evaluation: with zero online data, first build a golden set of 100 real human-organized questions, then automate gradually. These suggestions and the more complete landing framework are in RAG in Practice and Common Pitfalls and Anti-Patterns.
Beyond the checklist, there's a "soft" factor that determines success or failure: document governance. The same entity with multiple spellings ("OpenAI" and "openai," full company name vs. abbreviation), contradictory statements from different sources, frequently changing versions — all of these cause retrieval and generation to repeatedly fail. Governance investment doesn't directly show up in code, but it's the ceiling of retrieval quality. The relationship between document governance and RAG is like data cleaning to pretraining — methodology in Pretraining: Data and Objectives.
The #1 principle of RAG tuning
Fix retrieval first, then generation. 80% of RAG quality problems come from "the retrieved stuff is wrong," not "the model can't answer." Retrieval tuning methods (HyDE, reranking, hybrid retrieval) in RAG in Practice.
IX. Common Failure Mode Reference
| Failure Mode | Symptom | Root Cause | Countermeasure |
|---|---|---|---|
| Retrieval off-target | Answer unrelated to query | Query and doc semantics misaligned | Query rewriting, HyDE, hybrid retrieval |
| Context drowning | Relevant snippets drowned by noise | Chunk too large / top-k too high | Smaller chunks, reranking, compression |
| Low faithfulness | Answer fabricates details | Model free-play | Strong constraint prompts, citation format, evaluation fallback |
| Stale knowledge | Old answers | Index not updated | Incremental indexing, freshness filtering |
| Private data leak | Answer leaks unrelated docs | Missing permission filter | ACL filtering before retrieval |
A final piece of experience: RAG project failures, nine out of ten times the problem is data and retrieval, not the model — first get document governance and retrieval quality to 90 points, and the generation side naturally stabilizes.
X. RAG Advanced Topics and FAQ
1. Multi-hop Retrieval: What If the Answer Is in Two Places?
Single-turn retrieval can only get snippets "directly similar to the query," while many real problems need a reasoning chain:
Query: Where is the headquarters of that chip company acquired last week?
Hop 1: Retrieve "last week chip company acquisition" → find "XX company acquired by YY"
Hop 2: From XX company name, retrieve "XX company headquarters" → find the answerMulti-hop retrieval implementation: iterative retrieval (using previous answer as next query), query expansion (retrieve multiple sub-questions simultaneously), knowledge graph enhancement (using entity relation graphs to bridge cross-document associations). Higher complexity, but covers "indirect association" problems.
2. Caching and Incremental Updates
| Strategy | Approach | Benefit |
|---|---|---|
| Vector cache | Same/similar queries directly hit historical answers | Save 80%+ call cost |
| Incremental indexing | Document changes only rebuild changed chunks | Second-level updates |
| Dual-zone indexing | Hot zone (frequently asked) + cold zone (full) | Control cost |
| Freshness filtering | Filter by time/version at retrieval | Avoid old data answering |
How "alive" the knowledge base is often impacts RAG quality more than model selection.
3. Agentic RAG: Handing Retrieval to Agents
Traditional RAG is "one retrieval + one generation"; Agentic RAG lets LLM-Based Agents autonomously decide: which DB to query, how many times, is the result sufficient, should the strategy change. Typical forms:
- Routing: first judge question type, then choose the corresponding knowledge base;
- Tool-based: make retrieval, reranking, and web search tools, orchestrated by the Agent;
- Reflection: evaluate generation using self-evaluation, re-retrieve if insufficient (Self-RAG approach).
Agentic RAG is flexible but more expensive and less stable — suitable for "high-complexity problems needing multi-source info" scenarios.
4. Automating RAG Evaluation
Offline pipeline (run every change):
① Build golden set: 50–200 real questions + expected answers + expected citations
② Run retrieval: calculate Recall@k, MRR
③ Run generation: score faithfulness/relevance via RAGAS or LLM judge
④ Compare baseline (previous version) → regression report
⑤ Small-traffic online A/B validationMetric and tool details in Evaluations in Practice.
5. FAQ Quick Answers
| Question | Quick Answer |
|---|---|
| Can RAG eliminate hallucination? | Not completely, but can significantly reduce factual hallucination |
| Fine-tuning or RAG? | Knowledge → RAG, behavior → fine-tuning, can be combined |
| Can long context replace RAG? | No, cost and update speed don't allow it |
| How large a vector DB is needed? | 100K-level: pgvector/FAISS sufficient |
| Which embedding for Chinese? | bge / GTE / E5 Chinese versions |
| Where does RAG most commonly fail? | Retrieval quality, then chunk strategy |
Three-stage judgment for RAG
Naive can run, Advanced can improve scores, Modular/Agentic handles complex business. Don't jump to the heaviest architecture first — get index quality and evaluation sets right, and most projects already win half the battle.
Finally, viewing RAG in a larger context: it and "context engineering" are two sides of the same coin. Whether RAG, long context, or prompt compression, the essence is delivering "the most relevant information" to the model within limited inference budget. The only difference is "who does the filtering": RAG uses a retriever, long context lets the model search itself in large input, prompt compression uses a summarizer. Understanding this unified perspective lets you flexibly combine in specific scenarios: short docs → full text directly; long knowledge base → RAG; ultra-long single doc → RAG + long context hybrid. Tech buzzwords expire, but the proposition of "efficiently delivering relevant information to the model" does not.
One more note: the "division of labor" between retrieval and generation is fundamentally a choice between "externalized knowledge vs. internalized knowledge" — as long as knowledge changes, RAG deserves to remain the default option.
XI. RAG Variants and Adjacent Technologies
1. Main Variant Comparison
| Variant | Core Idea | Suitable For | Cost |
|---|---|---|---|
| Standard RAG | Vector retrieval + generation | General knowledge base | Retrieval quality sets ceiling |
| GraphRAG | Knowledge graph + graph traversal retrieval | Strongly associated, multi-hop problems | High graph construction cost |
| RAPTOR | Recursive clustering summary tree | Hierarchical understanding of long docs | Complex offline construction |
| HyDE | Generate hypothetical answer first, then retrieve | Zero-shot retrieval cold start | One extra generation step |
| RAG-Fusion | Multi-query + RRF fusion | Improve recall | Query cost doubles |
| Self-RAG | Self-reflection on whether to retrieve | Reduce unnecessary retrieval | Needs training reflection tokens |
2. RAG and Knowledge Graph QA
Knowledge graph QA (KGQA) answers "who, where, related to whom" problems with structured triples — high certainty but high construction and maintenance cost; RAG is flexible but lower certainty. In practice, "graph + RAG" hybrid is common: use the graph to bridge associations, use vector retrieval for semantics.
3. RAG + Fine-tuning Combination Patterns
| Combination | Solves What | How to Stack |
|---|---|---|
| RAG gives knowledge + fine-tuning gives style | Knowledge + customer service tone/format | Retrieval snippets + LoRA model |
| Fine-tune retriever | Poor retrieval quality | Fine-tune bi-encoder/reranker with annotated pairs |
| Fine-tune generator for citation format | Messy citation format | Fine-tune with citation-format instruction data |
Combination principle: fine-tuning doesn't change knowledge (that's RAG's job), RAG doesn't change style (that's fine-tuning's job).
4. RAG and MCP / Toolification
RAG retrieval itself can be made into a tool: expose "retrieve enterprise knowledge base" via MCP, letting LLM-Based Agents call it on demand during tasks — the standard form of Agentic RAG.
A prompt for researchers: RAG's open problems concentrate in three areas — "when to retrieve" (query intent judgment), "when is retrieval enough" (relevance measurement and stopping conditions), and "how to pass retrieval result credibility to the generation side" (citation and uncertainty propagation). These three correspond to Modular RAG's routing, reflection, and generation control modules, and are the core of Self-RAG, CRAG, and related work. Understanding these three open questions grasps RAG's research thread more than chasing new frameworks.
XII. RAG Key Paper Checklist
| Paper / Work | Year | One-Sentence |
|---|---|---|
| Retrieval-Augmented Generation (RAG) | 2020 | Founding work |
| DPR | 2020 | Dense retrieval bi-encoder |
| HyDE | 2022 | Zero-shot dense retrieval improvement |
| Self-RAG | 2023 | Self-reflective retrieval |
| CRAG | 2024 | Corrective retrieval quality |
| RAGAS | 2023 | Automated evaluation framework |
| GraphRAG | 2024 | Graph-augmented retrieval |
For newcomers, the recommended paper reading order: first read the 2020 RAG original paper to build the overall "retrieval-generation" picture, then read DPR to understand dense retrieval, then look at HyDE, Self-RAG, CRAG for evolution logic, and finally use RAGAS to understand evaluation. For each, just focus on "what problem was solved, what mechanism was introduced" — no need to derive every line. Paper deep-dive methods in Core Paper Deep Dives.
One sentence for RAG tech selection
The simpler the need, the more basic the solution: 80% of Q&A can use "standard RAG + reranking + evaluation loop"; GraphRAG and Agentic RAG are only worth introducing when multi-hop or autonomous decision-making is explicitly needed.
XIII. Further Reading
- Hallucination: Causes and Mitigation — RAG's mechanism for mitigating hallucination
- Context and Long Context — the tradeoff between long context and RAG
- Fine-tuning: SFT and Parameter-Efficient Fine-tuning — RAG vs. fine-tuning selection
- RAG in Practice — end-to-end building and tuning
- Evaluations in Practice — RAGAS and golden set
- Datasets and Benchmarks Archive — evaluation benchmarks and corpus resources
- LLM-Based Agents — RAG as the Agent's "knowledge tool"
References
- Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020) — RAG original paper (arXiv)
- Karpukhin et al. Dense Passage Retrieval for Open-Domain Question Answering (DPR, 2020) — Dense retrieval bi-encoder paper (arXiv)
- Gao et al. Precise Zero-Shot Dense Retrieval without Relevance Labels (HyDE, 2022) — HyDE paper (arXiv)
- Asai et al. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection (2023) — Self-RAG paper (arXiv)
- Yan et al. Corrective Retrieval Augmented Generation (CRAG, 2024) — CRAG paper (arXiv)
- Es et al. RAGAS: Automated Evaluation of Retrieval Augmented Generation (2023) — RAGAS evaluation framework paper (arXiv)
- Microsoft. Reinventing search with a new AI-powered Microsoft Bing (2023.2) — Bing Chat official blog
- Arcee-AI. Bespoke-Minerva-7B (Hugging Face) — RAG-specialized fine-tuned model weight page