Skip to content

RAG in Practice

At a glance Build a production-ready RAG system with an end-to-end pipeline of "document parsing → chunking → embedding → vector DB → hybrid search → reranking → generation → evaluation," and turn retrieval quality into a measurable, optimizable engineering discipline using chunk experiments, HyDE, failure mode tables, and evaluation metrics.

RAG in Practice ​

RAG (Retrieval-Augmented Generation) is fundamentally about moving "memory" out of model parameters and into an external index: the model no longer needs to memorize your private documents — it just needs to learn "retrieve first, then answer."

RAG is by far the most-used pattern in production-grade LLM applications. It simultaneously solves four problems: knowledge cutoff, hallucination, private data, and traceability. But RAG also has the worst tendency of "works in demos, breaks in production" — because the quality bottleneck is in the retrieval pipeline, not the model. This article gives an end-to-end deployable process with parameter suggestions and verification methods for each step. See Case Study: RAG for concepts and evolution background.

One, End-to-End Overview ​

The complete chain of a RAG system:

Offline (Index Construction)
Document → Parse → Clean → Chunking → Embedding → Write to vector DB (with metadata)
    │
    ▼
Online (Query)
User question → Query rewrite → Retrieve (vector + BM25) → Fuse → Rerank → Compose prompt → Generate → Validate output

We expand along the chain in order. Remember this master rule first: every ring's errors amplify downstream, and evaluation only happens at the end — so every minute you spend optimizing "retrieval recall" will ultimately reflect in answer quality.

The iteration rhythm for this chain:

StageKey MetricsIteration Rhythm
InitialEnd-to-end runsOne version is fine, don't linger on details
Retrieval optimizationRecall@k / Hit@kRe-run evaluation after every chunk/retrieval/rerank change
Generation optimizationFaithfulness / Answer relevanceRegression after prompt changes
ProductionOnline feedback + spot checksContinuous, feeding back into golden set

Two, Document Parsing & Chunking ​

1. Parsing: Turn "Files" into "Clean Text" ​

Document TypeParsing ToolNotes
PDF (text-based)pypdf, pdfplumberTables and multi-column layouts often break
PDF (scanned)OCR (Tesseract, PaddleOCR)OCR first, then parse — high cost
Word / PPTpython-docx, python-pptxWatch for headers, footers, comments
HTML / Web pagesBeautifulSoup, trafilaturaRemove nav, ads, scripts
Markdown / CodeDirect structural splitPreserve code block semantics

Parsing is RAG's #1 source of failures

Many "retrieval doesn't work" investigations ultimately discover that the parsing stage shredded PDF tables. Always spot-check parsed results manually before going live (5~10 pages per document type). Don't trust a parser's default output.

2. Chunking: The Art of Splitting ​

Chunks are the smallest unit of retrieval. Chunks that are too small lack context; chunks that are too large have diluted embeddings and imprecise recall. Two commonly used engineering approaches:

StrategyApproachProsCons
Fixed size + overlapCut 300~800 tokens per chunk, 10~20% overlapSimple, controllableMay break semantically complete sentences/paragraphs
Structure-aware (recursive)Split by paragraph → sentence → sentence priority (e.g., LangChain RecursiveCharacterTextSplitter)Respects text structureDepends on source document structure quality
Semantic splittingUse embedding similarity to find "semantic breakpoints"Consistent topic within each chunkHigh compute cost
python
# LangChain structure-aware splitting (the most commonly used default)
from langchain_text_splitters import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
    chunk_size=500,        # Target chars per chunk (estimate by token for mixed Chinese/English)
    chunk_overlap=80,      # Overlap chars between adjacent chunks, preventing key sentences from being cut at boundaries
    separators=["\n\n", "\n", "。", "!", "?", ". ", "! ", "? ", " ", ""],
)
chunks = splitter.split_text(long_document)
print(f"Created {len(chunks)} chunks, avg {sum(len(c) for c in chunks)/len(chunks):.0f} chars")

3. Chunk Size & Overlap: Let Experiments Speak ​

There is no "standard answer" for chunk size. You must experiment with your own documents and question set. The approach is to control variables:

python
import json

# Fixed set of 50~100 real questions (q, expected_chunk_id)
test_questions = json.load(open("rag_questions.json"))

def evaluate_chunk_size(chunk_size, chunk_overlap):
    """Rebuild index, measure Top-5 recall on these questions"""
    chunks = split_document(doc, chunk_size, chunk_overlap)
    index = build_vector_index(chunks)          # See Section Three
    hit = 0
    for q in test_questions:
        retrieved = search(index, q, k=5)        # See Section Four
        if q["expected_chunk_id"] in [r["id"] for r in retrieved]:
            hit += 1
    return hit / len(test_questions)

for size in [300, 500, 800, 1200]:
    print(size, evaluate_chunk_size(size, size // 5))

Default starting point and adjustment direction

Start with chunk_size ≈ 500 tokens, overlap ≈ 10~15%. If questions are "document-level overviews" (requiring global understanding), increase chunk size. If questions are "precise localization" (e.g., contract clauses), decrease chunk size and increase overlap. Always let recall rate / answer quality speak — don't rely on gut feeling.

Three, Embedding & Vector Database ​

1. Choosing an Embedding Model ​

Embedding maps text to dense vectors: semantically similar text has vectors that are closer in distance. Selection dimensions:

DimensionRecommendation
LanguageFor Chinese, prefer Chinese-optimized models (e.g., bge-large-zh, text-embedding-v3); for general English, use text-embedding-3-large or open-source E5/BGE series
Dimensions & cost1024~3072 dims are more accurate but more expensive in storage/retrieval; start with 768~1024 dims
Long-text supportWhen processing long document chunks, watch the model's max sequence length (most are 512~8192 tokens)
Open-source?For data staying within your intranet, choose open-source models (sentence-transformers ecosystem)
python
# Open-source embedding (one line with sentence-transformers)
pip install sentence-transformers

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("BAAI/bge-large-zh-v1.5")   # Chinese, 1024-dim
vec = model.encode("What does the third clause of the refund policy say?")             # Get vector
print(vec.shape)                                          # (1024,)

2. Vector Database Selection ​

Vector databases handle "storing vectors + approximate nearest neighbor search (ANN)." Decision table:

Vector DBFormScale LimitScenario
FAISSLibrary (embedded process)Tens of millionsPrototypes, offline batch, single machine
pgvectorPostgreSQL extensionMillions to tens of millionsExisting PG business systems, transactions + vectors coexist
Qdrant / MilvusStandalone serviceHundreds of millionsProduction-grade, multi-replica, hybrid search & filtering
Chroma / LanceDBLightweight embeddedMillionsLocal prototypes, teaching
python
# Minimal FAISS index (in-memory, good for prototyping)
pip install faiss-cpu

import faiss
import numpy as np

dim = 1024
index = faiss.IndexFlatIP(dim)        # Inner product index (works with normalized vectors = cosine similarity)
vecs = np.array([model.encode(c) for c in chunks])
faiss.normalize_L2(vecs)              # L2 normalize so inner product equals cosine
index.add(vecs)

q_vec = model.encode("Third clause of refund policy")
q_vec = q_vec / np.linalg.norm(q_vec)
D, I = index.search(q_vec[np.newaxis, :], k=5)
print("Top-5 chunk IDs:", I[0])       # Sorted by similarity

Limits of vector retrieval

Vector retrieval excels at "semantic similarity" but struggles with "exact keyword matching" (model numbers, IDs, proper nouns). This is precisely the gap that hybrid search in Section Four addresses.

Four, Retrieval: Hybrid Search, Query Rewriting & HyDE ​

1. Hybrid Search: BM25 + Vector ​

Keyword search (BM25) excels at exact terms; vector search excels at semantic near-synonyms. Production RAG almost always uses hybrid search with fusion:

python
# Illustration: fuse top-k results from BM25 and vector search by weight
# Rank fusion: for each result, score = sum( 1/(60 + rank_i) for each method i )

def reciprocal_rank_fusion(bm25_results, vector_results, k=60):
    scores = {}
    for rank, doc_id in enumerate(bm25_results + vector_results):
        scores[doc_id] = scores.get(doc_id, 0) + 1.0 / (k + rank + 1)
    return sorted(scores.items(), key=lambda x: -x[1])

Fusion typically yields higher recall than either alone. BM25 implementations: rank_bm25 library, Elasticsearch/OpenSearch, or Qdrant/Milvus built-in sparse retrieval.

2. Query Rewriting ​

A user's original question is often unsuitable for direct retrieval (too colloquial, too short, missing entities). Three common rewrite types:

Rewrite TypeExamplePurpose
Completion"What about refunds at that shop?" → "What is XX shop's refund policy?"Fill in missing entities
Multi-path rewriteOne sentence → 2~3 retrieval queriesExpand recall surface
Decomposition"Compare differences between A and B" → two sub-queriesDivide and conquer
python
# Query rewriting: minimal implementation (LLM generates multi-path queries)
def rewrite_queries(question: str) -> list[str]:
    prompt = f"""Rewrite the following question into 2~3 queries more suitable for retrieval.
Requirements: fill in missing entities, convert colloquial to formal, cover different expression angles. Only output queries, one per line.
Question: {question}"""
    text = call_llm(prompt)
    return [line.strip("- ").strip() for line in text.splitlines() if line.strip()]

# At retrieval time: retrieve with original question + rewritten queries, then fuse and deduplicate
all_results = []
for q in [question] + rewrite_queries(question):
    all_results += search(index, q, k=10)   # search uses vector/BM25 from above
# Then go through reciprocal rank fusion (see hybrid search) to merge and deduplicate

3. HyDE: Retrieving with "Hypothetical Answers" ​

HyDE (Hypothetical Document Embeddings) works on this idea: first have the LLM write a "hypothetical answer" out of thin air based on the question, then use that text's vector for retrieval. Because "answer-shaped text" sits closer to real documents in vector space than "question-shaped text," recall improves significantly (References).

python
def hyde_rewrite(question: str) -> str:
    """Have the LLM generate a hypothetical answer to use as retrieval query text"""
    prompt = f"""Based on the following question, write an approximate 100-word hypothetical answer.
Even if you're unsure of the facts, write something fluent and structurally complete that "looks like an encyclopedia entry."
Question: {question}
Hypothetical answer:"""
    return call_llm(prompt)   # Return this text, encode it, then retrieve

q_vec = model.encode(hyde_rewrite(question))   # Retrieve using HyDE vector

HyDE's cost and risk

HyDE adds one LLM call per query (latency + cost), and on tasks where "model hallucinations mislead retrieval" it can backfire. Try rewrites that don't incur extra call costs first, then decide whether to adopt HyDE.

Five, Reranking ​

Retrieve Top-50, then use a reranking model (cross-encoder) to fine-rank and take Top-5. A cross-encoder scores "query + document chunk" in pairs, with much higher accuracy than vector similarity — currently one of the best cost-to-accuracy levers in RAG:

python
from sentence_transformers import CrossEncoder

reranker = CrossEncoder("BAAI/bge-reranker-base")      # Chinese reranking model
pairs = [(question, chunk) for chunk in candidate_chunks]
scores = reranker.predict(pairs)                        # Score in pairs
top5 = [c for _, c in sorted(zip(scores, candidate_chunks), reverse=True)[:5]]
StageRetrieval TypeInputAccuracySpeed
Recall (Retrieve)Bi-encoder vector / BM25Single-sided encodingMediumExtremely fast (milliseconds)
Fine-rank (Rerank)Cross-encoderQuery + document pairHighSlow (only on candidate set)

Six, Generation: Composing the Prompt ​

Retrieved content must be composed into a "answerable" prompt. Three well-validated disciplines in the generation stage:

text
You are an internal knowledge base assistant. Answer only based on the provided materials.
Rules:
1. If materials are insufficient, clearly say "no relevant information in the materials" — do not fabricate
2. Cite source numbers, e.g. [1][2]
3. Answers must use only material content, no outside scope

【Materials】
[1] (Source: Sales Manual v3, page 5) ...
[2] (Source: Refund Policy 2025, clause 2) ...

【Question】
Customer asks: How long does a refund take to process?

【Answer】
  • "Answer only based on materials" is the first line of defense against hallucination (mechanism analysis in Hallucination: Causes & Mitigation).
  • Number materials and force citations: Answers become traceable, and downstream validation is easier.
  • "Say you don't know if materials are insufficient": Makes "I don't know" a valid output rather than forcing the model to make things up.

Don't force-stuff context

Slapping all retrieved results (even 20 chunks) into the prompt dilutes attention and wastes tokens. Only feed Top-3~5 chunks after reranking; the tradeoff with long context is covered in Context Windows & Long Text.

Seven, RAG Evaluation: Three Dimensions, All Essential ​

RAG output quality is jointly determined by retrieval quality and generation quality — evaluation must be layered:

DimensionMetricMeaningWhat's Evaluated
Retrieval recallRecall@k / Hit@kAre relevant documents in top-k?Retrieval pipeline
FaithfulnessCan every part of the answer be traced to materials?Anti-hallucinationGeneration pipeline
Answer relevanceDoes the answer address the question without going off-topic?End-to-endOverall system
python
# Quick quantification with RAGAS (faithfulness/relevance need LLM scoring)
pip install ragas

from ragas.metrics import faithfulness, answer_relevancy
from ragas import evaluate
# Build dataset (question / answer / contexts / ground_truth)
result = evaluate(dataset, metrics=[faithfulness, answer_relevancy])
print(result)   # Outputs scores 0~1 for each metric

See Evaluations in Practice for the full methodology of evaluation systems (golden sets, regression, LLM-as-a-judge bias).

Eight, Building a RAG Golden Set: Prerequisite for Evaluation ​

A prerequisite for RAG evaluation is a test set with ground-truth answers (question + expected answer). Key construction points:

PointDescription
Question sourceReal online logs, customer service records, not self-constructed
Answer formatChoose one: (a) expected chunk IDs to hit; (b) expected answer text
Difficulty layersSimple (directly searchable) / Medium (requires merging chunks) / Hard (requires cross-document reasoning)
QuantityStart with 50~100, add continuously

Once the golden set is built, re-run it after every chunk / retrieval / rerank / prompt change — this is RAG's regression testing. Full methods (layered evaluation, judge, cost control) are in Evaluations in Practice.

Nine, Common Failure Mode Table ​

FailureRoot CauseDebug & Fix
Retrieved content "looks related but isn't"Embedding domain driftSwitch to domain-finetuned embedding; add query rewriting; tune fusion weights
Exact IDs / model numbers not foundPure vector search weaker than keywordsAdd BM25 hybrid search; regex extraction + exact match when needed
Answers are irrelevantChunks too large / too fragmented, information splitTune chunk size; increase overlap; use reranking
Obvious fabrication (not in materials)Insufficient generation constraintsStrengthen "answer only based on materials"; add citation numbers; lower temperature
Old docs overshadow new onesIndex not versionedMetadata with version numbers; filter by latest version at retrieval
Bad at global questions about long docsSingle chunk lacks infoChunk summarization + secondary retrieval; or switch to long-context model
New documents don't take effect after uploadIndex not rebuilt / cachedBuild incremental index process driven by document changes

Ten, Toolchain Selection ​

LayerMain OptionsSelection Criteria
OrchestrationLangChain / LlamaIndex / customQuick prototype with frameworks, production consider custom or lightweight wrapper (see Framework & Tool Selection)
Vector DBFAISS / pgvector / Qdrant / MilvusSee Section Three decision table
EmbeddingOpenAI embeddings / BGE / E5Prefer Chinese models for Chinese scenarios
Rerankingbge-reranker / Cohere RerankAccuracy lever, strongly recommended to add
EvaluationRAGAS / custom golden setSee Section Seven

From minimal viable to production

A first version of RAG using 「PDF parsing + RecursiveCharacterTextSplitter + one open-source embedding + FAISS + hybrid search + bge-reranker + 3 prompt disciplines」can go live. Get that pipeline running and establish an evaluation set first, then swap in stronger components ring by ring — verify every upgrade with evaluation scores, not "it feels better." This "don't be hijacked by frameworks" principle also applies to Framework & Tool Selection.

Further Reading ​

References ​