Appearance
Build a RAG App from Scratch
Reading ten RAG tutorials teaches you less than actually running one RAG pipeline end to end. This article walks you through building a runnable personal knowledge base question-answering system in pure Python, from zero: no magic — you know what every line of code does. Once it runs, you'll genuinely understand what each step of RAG is there to solve.
Retrieval-Augmented Generation (RAG) is an architecture that lets a large language model (LLM) "look things up first, then answer": your private documents are split into chunks, embedded into vectors, and stored in an index; when a question arrives, the system first retrieves the most relevant chunks and hands them to the LLM together with the question to generate the answer. It targets the pain point that LLMs know nothing about your private material yet love to confabulate with a straight face — see the "hallucination" discussion in Large Language Models.
Why is it worth building once by hand? Because every stage of RAG carries its own trade-offs: how large a chunk can get before it loses meaning, what to do when retrieval misses, what to do when retrieval hits but the answer is still wrong. Frameworks (LangChain, LlamaIndex) wrap this pipeline into "a few lines of code", but for a beginner the wrapped-up parts are exactly where most of the traps hide. Write it by hand once, then use a framework, and you'll find you can actually read its error messages and understand its default behavior.
The full project breaks into five steps — here's the overall map first:
Personal documents (.md / .txt / PDF)
│
▼
① Document loading & chunking ──── answers "how knowledge becomes retrievable chunks"
▼
② Embedding & index building ───── answers "how chunks become comparable vectors"
▼
③ Retrieval (vector + BM25) ────── answers "how to fetch the most relevant chunks back"
▼
④ Prompt assembly & generation ─── answers "how to make the LLM answer only from the material"
▼
⑤ A complete runnable demo ─────── end-to-end run + tuning + evaluationPrerequisites
This article assumes you already have the basics: what RAG is (RAG Core Concepts), intuition for vectors and semantic search (Vector Databases and Semantic Search), and how to work with an LLM (Prompt Engineering). Fill whichever gap you have first, then come back to run the code. If you'd rather see what RAG looks like in production, read Perplexity and AI Search.
1. Tech Choices: Settle the Four Components First
RAG is assembled from four swappable components. The selection principle in one sentence: run whatever you can locally, and install as few services as possible. Pick "free + simple" at the prototype stage; swap by scale when you ship.
| Component | Options | Highlights | Best for |
|---|---|---|---|
| Embedding model | BGE (BAAI/bge-small-zh-v1.5, bge-m3) | Strong Chinese/multilingual quality, runs locally, MIT license | First choice for Chinese knowledge bases |
| OpenAI text-embedding-3-small / large | API-based, no local compute needed, multilingual | You already have an OpenAI API and want zero hassle | |
| M3E / Jina Embeddings / Cohere | Each has its own strengths (Chinese, multilingual, long text) | Pick from benchmark data | |
| Vector store | FAISS | In-process index, builds in seconds, no service process | Single-machine prototypes, small-to-mid scale (under a few million vectors) |
| Chroma | Embeds in Python, friendly API, supports persistence | Rapid prototyping, teaching demos | |
| pgvector | PostgreSQL extension, lives in the same database as your business data | Teams already on Postgres, need transactions and permissions | |
| LLM | OpenAI GPT series | Consistently good, zero fuss | Online services with budget to spare |
| Ollama local models (qwen2.5, llama3) | Free, private, fully under your control | Data must stay inside the intranet, offline scenarios | |
| DeepSeek / Tongyi / Doubao and other Chinese provider APIs | Strong Chinese, low prices | Chinese-language scenarios, cost-sensitive projects | |
| Framework | Hand-written (this article) | Every step transparent, easy to tune and to teach | Learning, fine-grained control |
| LlamaIndex | Wraps the whole "documents → knowledge base → Q&A" flow | Standing up knowledge-base QA quickly | |
| LangChain | Huge ecosystem, complete toolchain | Complex chains, multi-model integration |
Three judgment calls behind these choices:
- The embedding model sets the ceiling on retrieval quality. It maps your text into vectors — only then can similar documents land close together. For Chinese content use the BGE family (e.g.
BAAI/bge-small-zh-v1.5, 384 dimensions, about 400MB, runs on an ordinary laptop); consider lighter-weight models when lexical overlap is high and you want to save resources. For the model list and latest leaderboards, see Model & Leaderboard Cheat Sheet. - Don't agonize over the vector store. For a prototype, FAISS is enough — it's the retrieval layer's "dictionary implementation": accurate, fast, and zero operational burden. Migrate to pgvector or a dedicated service once you reach tens of millions of vectors or need filtering/updates/permissions.
- The LLM and the embedding model can come from different vendors. Computing embeddings locally with BGE while generating with GPT or a local Ollama model is perfectly fine — the two are entirely independent.
The honest truth about frameworks
Frameworks (LangChain / LlamaIndex) are not the trap; treating a framework as a "black box" is. This article has you write things by hand first because every RAG tuning lever lives in the details: where to place chunk boundaries, how to fuse retrieval results, how to write the prompt. Inside a framework these are "default parameters" you never see — and what you can't see, you can't debug. Run this article end to end first; picking up a framework afterwards will be much faster, because you know exactly what it's doing for you.
2. Environment Setup
bash
# Create and activate a virtual environment (never install dependencies into the system Python)
python -m venv .venv
source .venv/bin/activate # macOS / Linux
# .venv\Scripts\activate # Windows
# Core dependencies
pip install sentence-transformers faiss-cpu numpy openai jieba rank-bm25
# Optional: if you replace FAISS with Chroma
pip install chromadbVersion requirements: Python 3.10+, sentence-transformers 3.x, faiss-cpu 1.7+ (the CPU build is enough).
The first run downloads the embedding model (about 400MB); start with a smoke test:
python
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-small-zh-v1.5")
v = model.encode(["Hello, world"], normalize_embeddings=True)
print(v.shape) # (1, 384) — the model is ready3. Step-by-Step Implementation
Step 1: Document Loading and Chunking
Chunking cuts long documents into retrieval units. It is the first — and often the most important — tuning lever in RAG: chunks too small → each one carries too little information, so retrieval hits but answers come back incomplete; chunks too large → a single chunk mixes several topics, its vector gets diluted, and you waste context window. Common strategies:
| Strategy | Approach | Pros | Cons | Best for |
|---|---|---|---|---|
| Fixed window + overlap | Slice every N characters, with M characters of overlap between neighbors | Simplest to implement, most robust | May cut sentences or meaning in half | The general-purpose baseline — start here |
| Paragraph / heading split | Split on \n\n and heading levels | Keeps semantics intact | Long paragraphs can still exceed the window | Well-structured Markdown documents |
| Sentence aggregation | Split into sentences, then aggregate into chunks by semantics/length | Semantically complete boundaries | Depends on sentence-splitting quality | Q&A pairs, web page body text |
| Recursive character splitting | Fall back through separator levels, coarse to fine | Balances structure and length | Still needs size and overlap tuning | LangChain users |
Bottom line
Get it running with a "fixed 500-character window + overlap 50" first, then tune based on evaluation results. Don't chase fancy chunking strategies on day one of a project — most bad results come from later stages.
First, load the documents:
python
from pathlib import Path
def load_texts(directory: str = "./kb") -> list[dict]:
"""Load every .txt / .md file in a directory; returns [{"source": path, "text": full text}]."""
docs = []
for fp in sorted(Path(directory).glob("*")):
if fp.suffix.lower() not in {".txt", ".md"}:
continue
docs.append({"source": str(fp), "text": fp.read_text(encoding="utf-8")})
return docsNext, the chunking function. Key points: the fixed window does the heavy lifting, cuts are nudged toward natural boundaries such as "paragraphs and sentence ends" wherever possible, and neighboring chunks overlap so no sentence gets sliced in half:
python
def split_text(text: str, chunk_size: int = 500, overlap: int = 50) -> list[str]:
"""Fixed-window chunking + fallback to natural boundaries + overlap between neighbors.
overlap must be smaller than chunk_size, otherwise this loops forever.
"""
assert overlap < chunk_size, "overlap must be smaller than chunk_size"
if len(text) <= chunk_size:
return [text]
chunks, start = [], 0
while start < len(text):
end = min(start + chunk_size, len(text))
if end < len(text):
# Find the last natural boundary in the second half of the window (priority: paragraph > newline > sentence end)
for sep in ("\n\n", "\n", "。", ";"):
pos = text.rfind(sep, start + chunk_size // 2, end)
if pos != -1:
end = pos + len(sep)
break
chunks.append(text[start:end])
start = end - overlap
return chunks
# Organize into a uniform document structure
def build_chunks(docs: list[dict], chunk_size=500, overlap=50) -> list[dict]:
"""Returns [{"source", "chunk_index", "text"}]; indexing and retrieval below both work off this list."""
result = []
for doc in docs:
for i, chunk in enumerate(split_text(doc["text"], chunk_size, overlap)):
result.append({"source": doc["source"], "chunk_index": i, "text": chunk})
return resultStep 2: Embedding and Index Building
Embedding turns each chunk into a fixed-dimension vector; texts with similar semantics point in similar directions. Here we encode with the BGE Chinese model and normalize (after normalization, cosine similarity equals the vectors' inner product, so FAISS's inner-product index is all you need):
python
import json
import numpy as np
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-small-zh-v1.5") # 384 dims, Chinese/multilingual
chunks = build_chunks(load_texts())
texts = [c["text"] for c in chunks]
# normalize_embeddings=True: unit-norm vectors, so inner product == cosine similarity
vecs = model.encode(texts, normalize_embeddings=True)
print("chunks:", len(texts), "vector dim:", vecs.shape) # e.g. (N, 384)
np.save("kb_vecs.npy", vecs)
# Metadata (chunk text + source) is persisted separately, one-to-one with the vectors
json.dump(chunks, open("kb_meta.json", "w", encoding="utf-8"), ensure_ascii=False)Build the FAISS index. At prototype scale use IndexFlatIP (brute-force inner-product search — exact, but search cost grows linearly); switch to IndexIVFFlat (cluster-based coarse filtering) or an HNSW graph index once the data grows:
python
import faiss
dim = vecs.shape[1]
index = faiss.IndexFlatIP(dim) # inner-product index (vectors already normalized)
index.add(vecs.astype("float32")) # must be float32
faiss.write_index(index, "kb.index") # persist to disk; load and use later
print("index size:", index.ntotal)Where a vector database fits in this pipeline
You might ask: what does this have to do with a vector database? FAISS is merely an "in-process index" — no service, no network interface, no CRUD. You need a real vector database (Chroma, pgvector, or a cloud service) only when several clients query concurrently or you need insert/update/delete plus access control. FAISS for the prototype; a database when you productionize — don't let architectural overhead slow down the validation of your idea.
Step 3: Retrieval — Hybrid Vector + BM25
The goal of retrieval is to "recall the most relevant chunks". Vector-only recall has a known blind spot: queries whose wording overlaps lexically but whose meaning differs — search for "Apple's earnings report" and "apple" may float piles of fruit-related documents to the top. That's why the mature approach is hybrid retrieval with vectors + BM25, fused and re-ordered with RRF (Reciprocal Rank Fusion).
Vector recall:
python
def vector_search(query: str, k: int = 5) -> list[int]:
"""Return the chunk-list indices of the top-k hits."""
qv = model.encode([query], normalize_embeddings=True)
scores, ids = index.search(qv.astype("float32"), k)
return [int(i) for i in ids[0]]BM25 is the classic sparse retriever built on term frequency / inverse document frequency, and it is razor-sharp on exact keywords (tokenize Chinese text first):
python
import jieba
from rank_bm25 import BM25Okapi
tokenized = [list(jieba.cut(t)) for t in texts]
bm25 = BM25Okapi(tokenized)
def bm25_search(query: str, k: int = 5) -> list[int]:
scores = bm25.get_scores(list(jieba.cut(query)))
return scores.argsort()[::-1][:k].tolist()RRF fusion: ignore the absolute scores and look only at "ranks", merging the ordering information from the two result lists (the hyperparameter 60 is the conventional constant from the RRF paper):
python
def rrf_fuse(rank_lists: list[list[int]], k: int = 60, top_n: int = 5) -> list[int]:
"""Reciprocal Rank Fusion: score each ranked list by 1/(k+rank), then sum across lists."""
scores: dict[int, float] = {}
for lst in rank_lists:
for rank, doc_id in enumerate(lst):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)[:top_n]
def hybrid_search(query: str, top_n: int = 5) -> list[dict]:
vec_ids = vector_search(query, 10) # over-recall in each channel first, then fuse
bm25_ids = bm25_search(query, 10)
fused = rrf_fuse([vec_ids, bm25_ids], top_n=top_n)
return [chunks[i] for i in fused]Bottom line
Hybrid retrieval almost always beats a single channel. BM25 guards exact keywords, vectors guard semantic relevance, and RRF fusion makes the two complement each other. It's the first free upgrade on the road from "demo" to "usable".
Step 4: Generation — Prompt Assembly
The retrieved chunks then get "poured into" the prompt. Assembly follows three rules (see Prompt Engineering for details):
- Constrain the knowledge source: answer only from the provided material, and say "I don't know" outright when the material lacks the answer — this is the first gate against hallucination.
- Carry the sources: number each chunk and note its source file so every answer is "traceable".
- Control the context size: truncate each chunk to a sensible length so you neither blow the context window nor dilute attention with irrelevant text.
python
def build_prompt(query: str, hits: list[dict], max_chars: int = 800) -> str:
context = "\n\n".join(
f"[Chunk {i + 1}] (source: {h['source']})\n{h['text'][:max_chars]}"
for i, h in enumerate(hits)
)
return (
"You are a personal knowledge base assistant. Answer the question using only the material provided below; "
"if the material does not contain the answer, reply exactly 'The knowledge base contains no relevant information' and do not invent one.\n\n"
f"### Material\n{context}\n\n### Question\n{query}\n\n### Answer"
)
def generate(query: str, hits: list[dict]) -> str:
prompt = build_prompt(query, hits)
resp = client.chat.completions.create(
model=LLM_MODEL, messages=[{"role": "user", "content": prompt}],
temperature=0.3, # keep the temperature low for knowledge QA to curb improvisation
)
return resp.choices[0].message.contenttemperature=0.3 is the common setting for knowledge QA — you don't need the model to be "creative", you need it to be "faithful". For more techniques on controlling answer style, see The Prompt Playbook.
Step 5: A Complete, Runnable Demo
Wire the first four steps into one self-contained script, rag_demo.py. Usage: drop your personal documents into the kb/ directory, run once with --rebuild to build the index, then just ask:
python
"""
rag_demo.py — personal knowledge base Q&A (minimal pure-Python implementation)
Usage:
python rag_demo.py --rebuild # first run: build the index
python rag_demo.py "What is a vector database?" # ask a question
Dependencies: pip install sentence-transformers faiss-cpu numpy openai jieba rank-bm25
"""
import json
import sys
from pathlib import Path
import faiss
import jieba
import numpy as np
from openai import OpenAI
from rank_bm25 import BM25Okapi
from sentence_transformers import SentenceTransformer
# ── Configuration ──────────────────────────────────────
KB_DIR = Path("./kb") # put your .md / .txt files here
EMBED_MODEL = "BAAI/bge-small-zh-v1.5" # local embedding model
LLM_MODEL = "gpt-4o-mini" # generation model (local Ollama variant at the end)
CHUNK_SIZE, CHUNK_OVERLAP = 500, 50
TOP_N = 4
# ─────────────────────────────────────────────────────────
client = OpenAI() # requires the OPENAI_API_KEY environment variable
model = SentenceTransformer(EMBED_MODEL)
def load_texts() -> list[dict]:
docs = []
for fp in sorted(KB_DIR.glob("*")):
if fp.suffix.lower() in {".txt", ".md"}:
docs.append({"source": str(fp), "text": fp.read_text(encoding="utf-8")})
return docs
def split_text(text: str, chunk_size=CHUNK_SIZE, overlap=CHUNK_OVERLAP) -> list[str]:
assert overlap < chunk_size
if len(text) <= chunk_size:
return [text]
chunks, start = [], 0
while start < len(text):
end = min(start + chunk_size, len(text))
if end < len(text):
for sep in ("\n\n", "\n", "。", ";"):
pos = text.rfind(sep, start + chunk_size // 2, end)
if pos != -1:
end = pos + len(sep)
break
chunks.append(text[start:end])
start = end - overlap
return chunks
def build_index():
"""Chunk → embed → build the FAISS index → persist (metadata + vectors + index)."""
chunks = [
{"source": d["source"], "chunk_index": i, "text": c}
for d in load_texts()
for i, c in enumerate(split_text(d["text"]))
]
vecs = model.encode([c["text"] for c in chunks], normalize_embeddings=True)
index = faiss.IndexFlatIP(vecs.shape[1])
index.add(vecs.astype("float32"))
faiss.write_index(index, "kb.index")
np.save("kb_vecs.npy", vecs)
json.dump(chunks, open("kb_meta.json", "w", encoding="utf-8"), ensure_ascii=False)
print(f"Indexing complete: {len(chunks)} chunks")
def load_index():
global chunks, index, bm25
chunks = json.load(open("kb_meta.json", encoding="utf-8"))
index = faiss.read_index("kb.index")
vecs = np.load("kb_vecs.npy")
bm25 = BM25Okapi([list(jieba.cut(c["text"])) for c in chunks])
def vector_search(q: str, k: int) -> list[int]:
qv = model.encode([q], normalize_embeddings=True)
_, ids = index.search(qv.astype("float32"), k)
return [int(i) for i in ids[0]]
def bm25_search(q: str, k: int) -> list[int]:
scores = bm25.get_scores(list(jieba.cut(q)))
return scores.argsort()[::-1][:k].tolist()
def hybrid_search(q: str, top_n: int = TOP_N) -> list[dict]:
scores: dict[int, float] = {}
for lst in (vector_search(q, 10), bm25_search(q, 10)):
for rank, did in enumerate(lst):
scores[did] = scores.get(did, 0.0) + 1.0 / (60 + rank)
return [chunks[i] for i in sorted(scores, key=scores.get, reverse=True)[:top_n]]
def ask(query: str) -> str:
hits = hybrid_search(query)
context = "\n\n".join(
f"[Chunk {i + 1}] (source: {h['source']})\n{h['text'][:800]}"
for i, h in enumerate(hits)
)
prompt = (
"You are a personal knowledge base assistant. Answer the question using only the material below; "
"if the material contains no answer, reply exactly 'The knowledge base contains no relevant information' and do not invent one.\n\n"
f"### Material\n{context}\n\n### Question\n{query}\n\n### Answer"
)
resp = client.chat.completions.create(
model=LLM_MODEL, messages=[{"role": "user", "content": prompt}],
temperature=0.3,
)
answer = resp.choices[0].message.content
# Traceable citations: append the hit chunks to the answer for manual verification
refs = "\n".join(f"- {h['source']}" for h in hits)
return f"{answer}\n\n**Sources**:\n{refs}"
if __name__ == "__main__":
if "--rebuild" in sys.argv:
build_index()
load_index()
query = sys.argv[-1]
if query in {"--rebuild", "rag_demo.py"}:
print("Usage: python rag_demo.py \"your question\"")
else:
print(ask(query))Local Ollama variant: swap client and LLM_MODEL for the two lines below, leave every other line untouched, and the whole thing runs fully offline:
python
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
LLM_MODEL = "qwen2.5:7b" # any local model that supports tool calling; run ollama pull qwen2.5 firstOnce it runs you should observe two things: answers carry their sources, and the system can say "the knowledge base contains no relevant information" — exactly the two core improvements RAG offers over a bare LLM.
4. Key Tuning Levers: From "It Runs" to "It Works Well"
Almost every RAG problem comes down to "retrieval didn't find the right thing" or "generation didn't use it well". Tune in the order below — the returns diminish as you go:
| Tuning lever | How to tune | Rule-of-thumb values | Signal to watch |
|---|---|---|---|
| Chunk size | Experiment: run one evaluation pass each at 256 / 500 / 800 / 1200 | Start at 300–800 characters | Answers missing half the content → too small; answers drifting off topic → too large |
| Overlap | Adjust in step with chunk size | 10%–20% of the chunk size | Sentences cut off mid-way → increase |
| top-k | Increase the retrieval count and observe | 3–8 (recall 10 per channel before fusion) | Hits but incomplete answers → increase; context too noisy → decrease |
| Re-ranking (rerank) | Vector recall top 50 → cross-encoder rerank → keep the top 5 | Always do it (the payoff is large) | The right answer keeps landing in the runner-up slot |
| Source citations | Return chunk numbers and source files | Attach sources after each answer | Users can't verify the answer → mandatory |
Re-ranking is the highest-value step. Vector retrieval is "approximate matching"; a cross-encoder reranker (such as BAAI/bge-reranker-v2-m3) scores each "query + candidate chunk" pair through the model as a whole — far more accurate but slower, so "coarse recall first, precise rerank second" is the standard play:
python
from sentence_transformers import CrossEncoder
reranker = CrossEncoder("BAAI/bge-reranker-v2-m3")
cands = vector_search(query, 50) # coarse recall first
pairs = [(query, chunks[i]["text"]) for i in cands]
scores = reranker.predict(pairs)
top = [cands[i] for i in scores.argsort()[::-1][:TOP_N]] # then precise re-ranking5. Evaluation: Don't Trust "It Feels Good"
Evaluate RAG in two segments: the retrieval segment (did it find the right things?) and the generation segment (did it answer correctly?). Without evaluation you will never know whether a chunk tweak made things better or worse — this is the dividing line between dabbling and engineering.
Retrieval metrics (they require labeling "which chunks are correct for each question" by hand, or generating those labels automatically from document structure):
| Metric | Formula | Meaning |
|---|---|---|
| Recall@k | correct chunks hit / total correct chunks | How many of the correct chunks were recovered among the top k |
| MRR | Reciprocal of the first correct result's rank, averaged | How smoothly "the first hit is the right one" |
| Hit@k | Whether at least one correct chunk appears in the top k | Whether it's good enough to work with |
Generation metrics: faithfulness (can the answer be supported by the retrieved chunks?) and answer relevance (does it address the actual question?). The usual approach is LLM-as-judge, letting a second model do the scoring — see Building an LLM Evaluation Suite for a concrete build guide, and LLM Evaluation and Benchmarks for the full theory behind the metric system.
Evaluation discipline
- Build the eval set before you tune: collect 30–100 items of "question + correct chunks + reference answer". Every tuning decision made without an eval set is self-consolation.
- Measure retrieval and generation separately: however good generation is, it's wasted if retrieval recall is zero; optimize the retrieval metrics on their own first, then move on to generation.
- Change one variable, re-run: tweak chunk size, top-k, and rerank all at once and a good result tells you nothing about which change earned it.
6. Common Pitfalls
| Pitfall | Symptom | Root cause | Fix |
|---|---|---|---|
| Chunking breaks semantics | Answers missing half the content, or contradicting themselves | The fixed window lands right in the middle of a paragraph or sentence | Add overlap, fall back to natural boundaries, split on paragraphs |
| Retrieval misses → hallucination | The model fabricates with a straight face | Relevant chunks were never recalled, so the model has to wing it | Hybrid retrieval, larger recall, rerank, and a prompt that explicitly orders "say so when you don't know" |
| Unfiltered boilerplate | Retrieval keeps hitting footers / copyright notices / tables of contents | Footer text gets mixed into chunks and indexed | Strip boilerplate during cleaning and add filtering rules |
| Embeddings not normalized | Similarity scores aren't comparable | Cosine similarity and inner product used interchangeably | Standardize on normalize_embeddings=True |
| Context truncated | Answers stop halfway, or later material gets ignored | Chunk count × chunk length exceeds the window | Truncate individual chunks, control top-k, or move to a long-context model as needed |
| Index never refreshed | New documents can't be found | Documents were edited but the index was never rebuilt | Fold index building into the document update workflow (CI or a scheduled job) |
| Testing only one or two questions | "Feels" great, falls apart in production | Too small a sample that all happens to hit | Build a 30+ item eval set and track quantitative metrics |
| Retrieval hits but the answer is wrong | The material is right there, yet the answer is wrong | An unconstrained prompt, or distractors mixed into the chunks | Check the prompt constraints, check chunk purity, tune rerank |
For more general engineering anti-patterns, see Common Pitfalls and Anti-Patterns.
7. Where to Go Next
Once the basic RAG runs and has been evaluated, level up as needed:
- GraphRAG: multi-hop and global questions. Basic RAG is weak on "relational questions spanning multiple chunks" (such as "which documents all mention the same event"). Extracting the documents into an entity–relationship knowledge graph and retrieving over it supports multi-hop reasoning — see Knowledge Graphs and Knowledge Injection for the underlying principles.
- Agentic RAG: turn retrieval into a tool. Let the agent decide on its own "when to retrieve, how many times, whether to ask a clarifying question, whether to rewrite the query". Multi-turn clarification markedly improves vague questions, and query rewriting eases the mismatch between "how users ask" and "how documents are written" — this is the tool-calling loop from Building an Agent from Scratch applied to RAG; see Agent Core Concepts for the conceptual foundation.
- Streaming and caching. Stream generation output over SSE for a better experience; add semantic caching for high-frequency repeated questions (similar questions hit the cached answer directly), saving both tokens and latency — see Inference Optimization and Quantization and Deployment and Inference Optimization in Practice.
- Embedding fine-tuning. When domain jargon is dense and general-purpose embeddings can't tell things apart, fine-tune the embedding model on domain data — see Fine-Tuning and PEFT.
Closing verdict
A "passing" RAG system: stable metrics on an eval set, answers that carry citations, and an honest "I don't know" when the material doesn't cover it. Nail those three and it already outperforms most fake RAG setups of "bare LLM + a folder of knowledge files". Look up any unfamiliar term in the Glossary; for the big picture, see What Are the Hot AI Concepts.
Further Reading
- RAG Core Concepts — the full theory behind this article: RAG taxonomy, pipeline, weaknesses, and variants
- Vector Databases and Semantic Search — vector index internals, similarity measures, database selection
- Prompt Engineering — prompt assembly, context constraints, few-shot techniques
- The Prompt Playbook — the systematic methodology behind this article's prompts
- LLM Evaluation and Benchmarks — the theory behind retrieval and generation metrics
- Building an LLM Evaluation Suite — turning Section 5 into a runnable evaluation pipeline
- Building an Agent from Scratch — the tool-calling implementation behind Agentic RAG
- Knowledge Graphs and Knowledge Injection — the theoretical gateway to GraphRAG
- Perplexity and AI Search — a complete teardown of a production RAG product
- Common Pitfalls and Anti-Patterns — the complete 25-item version of Section 6
- Model & Leaderboard Cheat Sheet — up-to-date selection data for embeddings and LLMs
- Anatomy of the Overall Architecture — scaling this article's small project up to a production system
References
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (the original RAG paper, arXiv 2005.11401) — the paper that introduced the RAG architecture; essential reading
- Retrieval-Augmented Generation for Large Language Models: A Survey (arXiv 2312.10997) — a 2024 survey of RAG systems, covering the full chunking–retrieval–generation pipeline
- sentence-transformers documentation — official usage guide for embedding models such as BGE
- BAAI/bge model family (Hugging Face) — downloads and usage for the Chinese embedding models
- FAISS documentation — how to choose and tune vector index types
- Chroma documentation — getting started with the Python-embedded vector database
- pgvector documentation — the PostgreSQL vector extension
- rank-bm25 (PyPI) — a BM25 sparse retrieval implementation
- OpenAI Embeddings documentation — the cloud embedding API
- Ollama website — installation and model catalog for local LLMs