Theme
RAG in Practice
RAG (Retrieval-Augmented Generation) is fundamentally about moving "memory" out of model parameters and into an external index: the model no longer needs to memorize your private documents — it just needs to learn "retrieve first, then answer."
RAG is by far the most-used pattern in production-grade LLM applications. It simultaneously solves four problems: knowledge cutoff, hallucination, private data, and traceability. But RAG also has the worst tendency of "works in demos, breaks in production" — because the quality bottleneck is in the retrieval pipeline, not the model. This article gives an end-to-end deployable process with parameter suggestions and verification methods for each step. See Case Study: RAG for concepts and evolution background.
One, End-to-End Overview
The complete chain of a RAG system:
Offline (Index Construction)
Document → Parse → Clean → Chunking → Embedding → Write to vector DB (with metadata)
│
▼
Online (Query)
User question → Query rewrite → Retrieve (vector + BM25) → Fuse → Rerank → Compose prompt → Generate → Validate outputWe expand along the chain in order. Remember this master rule first: every ring's errors amplify downstream, and evaluation only happens at the end — so every minute you spend optimizing "retrieval recall" will ultimately reflect in answer quality.
The iteration rhythm for this chain:
| Stage | Key Metrics | Iteration Rhythm |
|---|---|---|
| Initial | End-to-end runs | One version is fine, don't linger on details |
| Retrieval optimization | Recall@k / Hit@k | Re-run evaluation after every chunk/retrieval/rerank change |
| Generation optimization | Faithfulness / Answer relevance | Regression after prompt changes |
| Production | Online feedback + spot checks | Continuous, feeding back into golden set |
Two, Document Parsing & Chunking
1. Parsing: Turn "Files" into "Clean Text"
| Document Type | Parsing Tool | Notes |
|---|---|---|
| PDF (text-based) | pypdf, pdfplumber | Tables and multi-column layouts often break |
| PDF (scanned) | OCR (Tesseract, PaddleOCR) | OCR first, then parse — high cost |
| Word / PPT | python-docx, python-pptx | Watch for headers, footers, comments |
| HTML / Web pages | BeautifulSoup, trafilatura | Remove nav, ads, scripts |
| Markdown / Code | Direct structural split | Preserve code block semantics |
Parsing is RAG's #1 source of failures
Many "retrieval doesn't work" investigations ultimately discover that the parsing stage shredded PDF tables. Always spot-check parsed results manually before going live (5~10 pages per document type). Don't trust a parser's default output.
2. Chunking: The Art of Splitting
Chunks are the smallest unit of retrieval. Chunks that are too small lack context; chunks that are too large have diluted embeddings and imprecise recall. Two commonly used engineering approaches:
| Strategy | Approach | Pros | Cons |
|---|---|---|---|
| Fixed size + overlap | Cut 300~800 tokens per chunk, 10~20% overlap | Simple, controllable | May break semantically complete sentences/paragraphs |
| Structure-aware (recursive) | Split by paragraph → sentence → sentence priority (e.g., LangChain RecursiveCharacterTextSplitter) | Respects text structure | Depends on source document structure quality |
| Semantic splitting | Use embedding similarity to find "semantic breakpoints" | Consistent topic within each chunk | High compute cost |
python
# LangChain structure-aware splitting (the most commonly used default)
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=500, # Target chars per chunk (estimate by token for mixed Chinese/English)
chunk_overlap=80, # Overlap chars between adjacent chunks, preventing key sentences from being cut at boundaries
separators=["\n\n", "\n", "。", "!", "?", ". ", "! ", "? ", " ", ""],
)
chunks = splitter.split_text(long_document)
print(f"Created {len(chunks)} chunks, avg {sum(len(c) for c in chunks)/len(chunks):.0f} chars")3. Chunk Size & Overlap: Let Experiments Speak
There is no "standard answer" for chunk size. You must experiment with your own documents and question set. The approach is to control variables:
python
import json
# Fixed set of 50~100 real questions (q, expected_chunk_id)
test_questions = json.load(open("rag_questions.json"))
def evaluate_chunk_size(chunk_size, chunk_overlap):
"""Rebuild index, measure Top-5 recall on these questions"""
chunks = split_document(doc, chunk_size, chunk_overlap)
index = build_vector_index(chunks) # See Section Three
hit = 0
for q in test_questions:
retrieved = search(index, q, k=5) # See Section Four
if q["expected_chunk_id"] in [r["id"] for r in retrieved]:
hit += 1
return hit / len(test_questions)
for size in [300, 500, 800, 1200]:
print(size, evaluate_chunk_size(size, size // 5))Default starting point and adjustment direction
Start with chunk_size ≈ 500 tokens, overlap ≈ 10~15%. If questions are "document-level overviews" (requiring global understanding), increase chunk size. If questions are "precise localization" (e.g., contract clauses), decrease chunk size and increase overlap. Always let recall rate / answer quality speak — don't rely on gut feeling.
Three, Embedding & Vector Database
1. Choosing an Embedding Model
Embedding maps text to dense vectors: semantically similar text has vectors that are closer in distance. Selection dimensions:
| Dimension | Recommendation |
|---|---|
| Language | For Chinese, prefer Chinese-optimized models (e.g., bge-large-zh, text-embedding-v3); for general English, use text-embedding-3-large or open-source E5/BGE series |
| Dimensions & cost | 1024~3072 dims are more accurate but more expensive in storage/retrieval; start with 768~1024 dims |
| Long-text support | When processing long document chunks, watch the model's max sequence length (most are 512~8192 tokens) |
| Open-source? | For data staying within your intranet, choose open-source models (sentence-transformers ecosystem) |
python
# Open-source embedding (one line with sentence-transformers)
pip install sentence-transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-large-zh-v1.5") # Chinese, 1024-dim
vec = model.encode("What does the third clause of the refund policy say?") # Get vector
print(vec.shape) # (1024,)2. Vector Database Selection
Vector databases handle "storing vectors + approximate nearest neighbor search (ANN)." Decision table:
| Vector DB | Form | Scale Limit | Scenario |
|---|---|---|---|
| FAISS | Library (embedded process) | Tens of millions | Prototypes, offline batch, single machine |
| pgvector | PostgreSQL extension | Millions to tens of millions | Existing PG business systems, transactions + vectors coexist |
| Qdrant / Milvus | Standalone service | Hundreds of millions | Production-grade, multi-replica, hybrid search & filtering |
| Chroma / LanceDB | Lightweight embedded | Millions | Local prototypes, teaching |
python
# Minimal FAISS index (in-memory, good for prototyping)
pip install faiss-cpu
import faiss
import numpy as np
dim = 1024
index = faiss.IndexFlatIP(dim) # Inner product index (works with normalized vectors = cosine similarity)
vecs = np.array([model.encode(c) for c in chunks])
faiss.normalize_L2(vecs) # L2 normalize so inner product equals cosine
index.add(vecs)
q_vec = model.encode("Third clause of refund policy")
q_vec = q_vec / np.linalg.norm(q_vec)
D, I = index.search(q_vec[np.newaxis, :], k=5)
print("Top-5 chunk IDs:", I[0]) # Sorted by similarityLimits of vector retrieval
Vector retrieval excels at "semantic similarity" but struggles with "exact keyword matching" (model numbers, IDs, proper nouns). This is precisely the gap that hybrid search in Section Four addresses.
Four, Retrieval: Hybrid Search, Query Rewriting & HyDE
1. Hybrid Search: BM25 + Vector
Keyword search (BM25) excels at exact terms; vector search excels at semantic near-synonyms. Production RAG almost always uses hybrid search with fusion:
python
# Illustration: fuse top-k results from BM25 and vector search by weight
# Rank fusion: for each result, score = sum( 1/(60 + rank_i) for each method i )
def reciprocal_rank_fusion(bm25_results, vector_results, k=60):
scores = {}
for rank, doc_id in enumerate(bm25_results + vector_results):
scores[doc_id] = scores.get(doc_id, 0) + 1.0 / (k + rank + 1)
return sorted(scores.items(), key=lambda x: -x[1])Fusion typically yields higher recall than either alone. BM25 implementations: rank_bm25 library, Elasticsearch/OpenSearch, or Qdrant/Milvus built-in sparse retrieval.
2. Query Rewriting
A user's original question is often unsuitable for direct retrieval (too colloquial, too short, missing entities). Three common rewrite types:
| Rewrite Type | Example | Purpose |
|---|---|---|
| Completion | "What about refunds at that shop?" → "What is XX shop's refund policy?" | Fill in missing entities |
| Multi-path rewrite | One sentence → 2~3 retrieval queries | Expand recall surface |
| Decomposition | "Compare differences between A and B" → two sub-queries | Divide and conquer |
python
# Query rewriting: minimal implementation (LLM generates multi-path queries)
def rewrite_queries(question: str) -> list[str]:
prompt = f"""Rewrite the following question into 2~3 queries more suitable for retrieval.
Requirements: fill in missing entities, convert colloquial to formal, cover different expression angles. Only output queries, one per line.
Question: {question}"""
text = call_llm(prompt)
return [line.strip("- ").strip() for line in text.splitlines() if line.strip()]
# At retrieval time: retrieve with original question + rewritten queries, then fuse and deduplicate
all_results = []
for q in [question] + rewrite_queries(question):
all_results += search(index, q, k=10) # search uses vector/BM25 from above
# Then go through reciprocal rank fusion (see hybrid search) to merge and deduplicate3. HyDE: Retrieving with "Hypothetical Answers"
HyDE (Hypothetical Document Embeddings) works on this idea: first have the LLM write a "hypothetical answer" out of thin air based on the question, then use that text's vector for retrieval. Because "answer-shaped text" sits closer to real documents in vector space than "question-shaped text," recall improves significantly (References).
python
def hyde_rewrite(question: str) -> str:
"""Have the LLM generate a hypothetical answer to use as retrieval query text"""
prompt = f"""Based on the following question, write an approximate 100-word hypothetical answer.
Even if you're unsure of the facts, write something fluent and structurally complete that "looks like an encyclopedia entry."
Question: {question}
Hypothetical answer:"""
return call_llm(prompt) # Return this text, encode it, then retrieve
q_vec = model.encode(hyde_rewrite(question)) # Retrieve using HyDE vectorHyDE's cost and risk
HyDE adds one LLM call per query (latency + cost), and on tasks where "model hallucinations mislead retrieval" it can backfire. Try rewrites that don't incur extra call costs first, then decide whether to adopt HyDE.
Five, Reranking
Retrieve Top-50, then use a reranking model (cross-encoder) to fine-rank and take Top-5. A cross-encoder scores "query + document chunk" in pairs, with much higher accuracy than vector similarity — currently one of the best cost-to-accuracy levers in RAG:
python
from sentence_transformers import CrossEncoder
reranker = CrossEncoder("BAAI/bge-reranker-base") # Chinese reranking model
pairs = [(question, chunk) for chunk in candidate_chunks]
scores = reranker.predict(pairs) # Score in pairs
top5 = [c for _, c in sorted(zip(scores, candidate_chunks), reverse=True)[:5]]| Stage | Retrieval Type | Input | Accuracy | Speed |
|---|---|---|---|---|
| Recall (Retrieve) | Bi-encoder vector / BM25 | Single-sided encoding | Medium | Extremely fast (milliseconds) |
| Fine-rank (Rerank) | Cross-encoder | Query + document pair | High | Slow (only on candidate set) |
Six, Generation: Composing the Prompt
Retrieved content must be composed into a "answerable" prompt. Three well-validated disciplines in the generation stage:
text
You are an internal knowledge base assistant. Answer only based on the provided materials.
Rules:
1. If materials are insufficient, clearly say "no relevant information in the materials" — do not fabricate
2. Cite source numbers, e.g. [1][2]
3. Answers must use only material content, no outside scope
【Materials】
[1] (Source: Sales Manual v3, page 5) ...
[2] (Source: Refund Policy 2025, clause 2) ...
【Question】
Customer asks: How long does a refund take to process?
【Answer】- "Answer only based on materials" is the first line of defense against hallucination (mechanism analysis in Hallucination: Causes & Mitigation).
- Number materials and force citations: Answers become traceable, and downstream validation is easier.
- "Say you don't know if materials are insufficient": Makes "I don't know" a valid output rather than forcing the model to make things up.
Don't force-stuff context
Slapping all retrieved results (even 20 chunks) into the prompt dilutes attention and wastes tokens. Only feed Top-3~5 chunks after reranking; the tradeoff with long context is covered in Context Windows & Long Text.
Seven, RAG Evaluation: Three Dimensions, All Essential
RAG output quality is jointly determined by retrieval quality and generation quality — evaluation must be layered:
| Dimension | Metric | Meaning | What's Evaluated |
|---|---|---|---|
| Retrieval recall | Recall@k / Hit@k | Are relevant documents in top-k? | Retrieval pipeline |
| Faithfulness | Can every part of the answer be traced to materials? | Anti-hallucination | Generation pipeline |
| Answer relevance | Does the answer address the question without going off-topic? | End-to-end | Overall system |
python
# Quick quantification with RAGAS (faithfulness/relevance need LLM scoring)
pip install ragas
from ragas.metrics import faithfulness, answer_relevancy
from ragas import evaluate
# Build dataset (question / answer / contexts / ground_truth)
result = evaluate(dataset, metrics=[faithfulness, answer_relevancy])
print(result) # Outputs scores 0~1 for each metricSee Evaluations in Practice for the full methodology of evaluation systems (golden sets, regression, LLM-as-a-judge bias).
Eight, Building a RAG Golden Set: Prerequisite for Evaluation
A prerequisite for RAG evaluation is a test set with ground-truth answers (question + expected answer). Key construction points:
| Point | Description |
|---|---|
| Question source | Real online logs, customer service records, not self-constructed |
| Answer format | Choose one: (a) expected chunk IDs to hit; (b) expected answer text |
| Difficulty layers | Simple (directly searchable) / Medium (requires merging chunks) / Hard (requires cross-document reasoning) |
| Quantity | Start with 50~100, add continuously |
Once the golden set is built, re-run it after every chunk / retrieval / rerank / prompt change — this is RAG's regression testing. Full methods (layered evaluation, judge, cost control) are in Evaluations in Practice.
Nine, Common Failure Mode Table
| Failure | Root Cause | Debug & Fix |
|---|---|---|
| Retrieved content "looks related but isn't" | Embedding domain drift | Switch to domain-finetuned embedding; add query rewriting; tune fusion weights |
| Exact IDs / model numbers not found | Pure vector search weaker than keywords | Add BM25 hybrid search; regex extraction + exact match when needed |
| Answers are irrelevant | Chunks too large / too fragmented, information split | Tune chunk size; increase overlap; use reranking |
| Obvious fabrication (not in materials) | Insufficient generation constraints | Strengthen "answer only based on materials"; add citation numbers; lower temperature |
| Old docs overshadow new ones | Index not versioned | Metadata with version numbers; filter by latest version at retrieval |
| Bad at global questions about long docs | Single chunk lacks info | Chunk summarization + secondary retrieval; or switch to long-context model |
| New documents don't take effect after upload | Index not rebuilt / cached | Build incremental index process driven by document changes |
Ten, Toolchain Selection
| Layer | Main Options | Selection Criteria |
|---|---|---|
| Orchestration | LangChain / LlamaIndex / custom | Quick prototype with frameworks, production consider custom or lightweight wrapper (see Framework & Tool Selection) |
| Vector DB | FAISS / pgvector / Qdrant / Milvus | See Section Three decision table |
| Embedding | OpenAI embeddings / BGE / E5 | Prefer Chinese models for Chinese scenarios |
| Reranking | bge-reranker / Cohere Rerank | Accuracy lever, strongly recommended to add |
| Evaluation | RAGAS / custom golden set | See Section Seven |
From minimal viable to production
A first version of RAG using 「PDF parsing + RecursiveCharacterTextSplitter + one open-source embedding + FAISS + hybrid search + bge-reranker + 3 prompt disciplines」can go live. Get that pipeline running and establish an evaluation set first, then swap in stronger components ring by ring — verify every upgrade with evaluation scores, not "it feels better." This "don't be hijacked by frameworks" principle also applies to Framework & Tool Selection.
Further Reading
- Case Study: RAG (Retrieval-Augmented Generation) —— RAG evolution (Naive → Advanced → Modular) and concept overview
- Hallucination: Causes & Mitigation —— The mechanism behind RAG anti-hallucination: why "answer only based on materials" works
- Evaluations in Practice —— Full implementation of golden sets, regression testing, LLM-as-a-judge
- Context Windows & Long Text —— RAG vs. long context tradeoffs
- Prompt Engineering in Practice —— Systematic methods for composing generation-stage prompts
- Framework & Tool Selection —— Complete decision tables for orchestration frameworks and vector databases
- Common Pitfalls and Anti-Patterns —— RAG-related pitfalls like "stuffing context to the brim"
References
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv:2005.11401) —— Original RAG paper, proposed by Facebook AI
- Precise Zero-Shot Dense Retrieval without Relevance Labels (HyDE, arXiv:2212.10496) —— Original HyDE paper
- MTEB: Massive Text Embedding Benchmark (arXiv:2210.07316) —— Retrieval/embedding model leaderboard
- BGE/FlagEmbedding (GitHub) —— Official repo for BGE series embeddings and reranking models
- FAISS (GitHub) —— Meta's open-source vector retrieval library
- Qdrant documentation —— Production-grade vector database docs
- pgvector (GitHub) —— PostgreSQL vector extension
- LlamaIndex documentation —— Data framework, including RAG components for each stage
- LangChain text splitter docs —— Official usage of RecursiveCharacterTextSplitter
- RAGAS (GitHub) —— Open-source RAG evaluation library