Skip to content

Build a RAG App from Scratch

At a glance Build a runnable personal-knowledge-base RAG system from scratch in Python — document loading and chunking, embedding-based indexing, hybrid vector and BM25 retrieval, and prompt assembly with generation — with fully runnable code plus tuning, evaluation, common pitfalls, and next steps.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Build a RAG App from Scratch ​

Reading ten RAG tutorials teaches you less than actually running one RAG pipeline end to end. This article walks you through building a runnable personal knowledge base question-answering system in pure Python, from zero: no magic — you know what every line of code does. Once it runs, you'll genuinely understand what each step of RAG is there to solve.

Retrieval-Augmented Generation (RAG) is an architecture that lets a large language model (LLM) "look things up first, then answer": your private documents are split into chunks, embedded into vectors, and stored in an index; when a question arrives, the system first retrieves the most relevant chunks and hands them to the LLM together with the question to generate the answer. It targets the pain point that LLMs know nothing about your private material yet love to confabulate with a straight face — see the "hallucination" discussion in Large Language Models.

Why is it worth building once by hand? Because every stage of RAG carries its own trade-offs: how large a chunk can get before it loses meaning, what to do when retrieval misses, what to do when retrieval hits but the answer is still wrong. Frameworks (LangChain, LlamaIndex) wrap this pipeline into "a few lines of code", but for a beginner the wrapped-up parts are exactly where most of the traps hide. Write it by hand once, then use a framework, and you'll find you can actually read its error messages and understand its default behavior.

The full project breaks into five steps — here's the overall map first:

Personal documents (.md / .txt / PDF)
   │
   ▼
① Document loading & chunking ──── answers "how knowledge becomes retrievable chunks"
   ▼
② Embedding & index building ───── answers "how chunks become comparable vectors"
   ▼
③ Retrieval (vector + BM25) ────── answers "how to fetch the most relevant chunks back"
   ▼
④ Prompt assembly & generation ─── answers "how to make the LLM answer only from the material"
   ▼
⑤ A complete runnable demo ─────── end-to-end run + tuning + evaluation

Prerequisites

This article assumes you already have the basics: what RAG is (RAG Core Concepts), intuition for vectors and semantic search (Vector Databases and Semantic Search), and how to work with an LLM (Prompt Engineering). Fill whichever gap you have first, then come back to run the code. If you'd rather see what RAG looks like in production, read Perplexity and AI Search.

1. Tech Choices: Settle the Four Components First ​

RAG is assembled from four swappable components. The selection principle in one sentence: run whatever you can locally, and install as few services as possible. Pick "free + simple" at the prototype stage; swap by scale when you ship.

ComponentOptionsHighlightsBest for
Embedding modelBGE (BAAI/bge-small-zh-v1.5, bge-m3)Strong Chinese/multilingual quality, runs locally, MIT licenseFirst choice for Chinese knowledge bases
OpenAI text-embedding-3-small / largeAPI-based, no local compute needed, multilingualYou already have an OpenAI API and want zero hassle
M3E / Jina Embeddings / CohereEach has its own strengths (Chinese, multilingual, long text)Pick from benchmark data
Vector storeFAISSIn-process index, builds in seconds, no service processSingle-machine prototypes, small-to-mid scale (under a few million vectors)
ChromaEmbeds in Python, friendly API, supports persistenceRapid prototyping, teaching demos
pgvectorPostgreSQL extension, lives in the same database as your business dataTeams already on Postgres, need transactions and permissions
LLMOpenAI GPT seriesConsistently good, zero fussOnline services with budget to spare
Ollama local models (qwen2.5, llama3)Free, private, fully under your controlData must stay inside the intranet, offline scenarios
DeepSeek / Tongyi / Doubao and other Chinese provider APIsStrong Chinese, low pricesChinese-language scenarios, cost-sensitive projects
FrameworkHand-written (this article)Every step transparent, easy to tune and to teachLearning, fine-grained control
LlamaIndexWraps the whole "documents → knowledge base → Q&A" flowStanding up knowledge-base QA quickly
LangChainHuge ecosystem, complete toolchainComplex chains, multi-model integration

Three judgment calls behind these choices:

  • The embedding model sets the ceiling on retrieval quality. It maps your text into vectors — only then can similar documents land close together. For Chinese content use the BGE family (e.g. BAAI/bge-small-zh-v1.5, 384 dimensions, about 400MB, runs on an ordinary laptop); consider lighter-weight models when lexical overlap is high and you want to save resources. For the model list and latest leaderboards, see Model & Leaderboard Cheat Sheet.
  • Don't agonize over the vector store. For a prototype, FAISS is enough — it's the retrieval layer's "dictionary implementation": accurate, fast, and zero operational burden. Migrate to pgvector or a dedicated service once you reach tens of millions of vectors or need filtering/updates/permissions.
  • The LLM and the embedding model can come from different vendors. Computing embeddings locally with BGE while generating with GPT or a local Ollama model is perfectly fine — the two are entirely independent.

The honest truth about frameworks

Frameworks (LangChain / LlamaIndex) are not the trap; treating a framework as a "black box" is. This article has you write things by hand first because every RAG tuning lever lives in the details: where to place chunk boundaries, how to fuse retrieval results, how to write the prompt. Inside a framework these are "default parameters" you never see — and what you can't see, you can't debug. Run this article end to end first; picking up a framework afterwards will be much faster, because you know exactly what it's doing for you.

2. Environment Setup ​

bash
# Create and activate a virtual environment (never install dependencies into the system Python)
python -m venv .venv
source .venv/bin/activate        # macOS / Linux
# .venv\Scripts\activate         # Windows

# Core dependencies
pip install sentence-transformers faiss-cpu numpy openai jieba rank-bm25
# Optional: if you replace FAISS with Chroma
pip install chromadb

Version requirements: Python 3.10+, sentence-transformers 3.x, faiss-cpu 1.7+ (the CPU build is enough).

The first run downloads the embedding model (about 400MB); start with a smoke test:

python
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-small-zh-v1.5")
v = model.encode(["Hello, world"], normalize_embeddings=True)
print(v.shape)   # (1, 384) — the model is ready

3. Step-by-Step Implementation ​

Step 1: Document Loading and Chunking ​

Chunking cuts long documents into retrieval units. It is the first — and often the most important — tuning lever in RAG: chunks too small → each one carries too little information, so retrieval hits but answers come back incomplete; chunks too large → a single chunk mixes several topics, its vector gets diluted, and you waste context window. Common strategies:

StrategyApproachProsConsBest for
Fixed window + overlapSlice every N characters, with M characters of overlap between neighborsSimplest to implement, most robustMay cut sentences or meaning in halfThe general-purpose baseline — start here
Paragraph / heading splitSplit on \n\n and heading levelsKeeps semantics intactLong paragraphs can still exceed the windowWell-structured Markdown documents
Sentence aggregationSplit into sentences, then aggregate into chunks by semantics/lengthSemantically complete boundariesDepends on sentence-splitting qualityQ&A pairs, web page body text
Recursive character splittingFall back through separator levels, coarse to fineBalances structure and lengthStill needs size and overlap tuningLangChain users

Bottom line

Get it running with a "fixed 500-character window + overlap 50" first, then tune based on evaluation results. Don't chase fancy chunking strategies on day one of a project — most bad results come from later stages.

First, load the documents:

python
from pathlib import Path

def load_texts(directory: str = "./kb") -> list[dict]:
    """Load every .txt / .md file in a directory; returns [{"source": path, "text": full text}]."""
    docs = []
    for fp in sorted(Path(directory).glob("*")):
        if fp.suffix.lower() not in {".txt", ".md"}:
            continue
        docs.append({"source": str(fp), "text": fp.read_text(encoding="utf-8")})
    return docs

Next, the chunking function. Key points: the fixed window does the heavy lifting, cuts are nudged toward natural boundaries such as "paragraphs and sentence ends" wherever possible, and neighboring chunks overlap so no sentence gets sliced in half:

python
def split_text(text: str, chunk_size: int = 500, overlap: int = 50) -> list[str]:
    """Fixed-window chunking + fallback to natural boundaries + overlap between neighbors.

    overlap must be smaller than chunk_size, otherwise this loops forever.
    """
    assert overlap < chunk_size, "overlap must be smaller than chunk_size"
    if len(text) <= chunk_size:
        return [text]

    chunks, start = [], 0
    while start < len(text):
        end = min(start + chunk_size, len(text))
        if end < len(text):
            # Find the last natural boundary in the second half of the window (priority: paragraph > newline > sentence end)
            for sep in ("\n\n", "\n", "。", ";"):
                pos = text.rfind(sep, start + chunk_size // 2, end)
                if pos != -1:
                    end = pos + len(sep)
                    break
        chunks.append(text[start:end])
        start = end - overlap
    return chunks

# Organize into a uniform document structure
def build_chunks(docs: list[dict], chunk_size=500, overlap=50) -> list[dict]:
    """Returns [{"source", "chunk_index", "text"}]; indexing and retrieval below both work off this list."""
    result = []
    for doc in docs:
        for i, chunk in enumerate(split_text(doc["text"], chunk_size, overlap)):
            result.append({"source": doc["source"], "chunk_index": i, "text": chunk})
    return result

Step 2: Embedding and Index Building ​

Embedding turns each chunk into a fixed-dimension vector; texts with similar semantics point in similar directions. Here we encode with the BGE Chinese model and normalize (after normalization, cosine similarity equals the vectors' inner product, so FAISS's inner-product index is all you need):

python
import json
import numpy as np
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("BAAI/bge-small-zh-v1.5")   # 384 dims, Chinese/multilingual

chunks = build_chunks(load_texts())
texts = [c["text"] for c in chunks]

# normalize_embeddings=True: unit-norm vectors, so inner product == cosine similarity
vecs = model.encode(texts, normalize_embeddings=True)
print("chunks:", len(texts), "vector dim:", vecs.shape)   # e.g. (N, 384)

np.save("kb_vecs.npy", vecs)

# Metadata (chunk text + source) is persisted separately, one-to-one with the vectors
json.dump(chunks, open("kb_meta.json", "w", encoding="utf-8"), ensure_ascii=False)

Build the FAISS index. At prototype scale use IndexFlatIP (brute-force inner-product search — exact, but search cost grows linearly); switch to IndexIVFFlat (cluster-based coarse filtering) or an HNSW graph index once the data grows:

python
import faiss

dim = vecs.shape[1]
index = faiss.IndexFlatIP(dim)                      # inner-product index (vectors already normalized)
index.add(vecs.astype("float32"))                   # must be float32
faiss.write_index(index, "kb.index")                # persist to disk; load and use later
print("index size:", index.ntotal)

Where a vector database fits in this pipeline

You might ask: what does this have to do with a vector database? FAISS is merely an "in-process index" — no service, no network interface, no CRUD. You need a real vector database (Chroma, pgvector, or a cloud service) only when several clients query concurrently or you need insert/update/delete plus access control. FAISS for the prototype; a database when you productionize — don't let architectural overhead slow down the validation of your idea.

Step 3: Retrieval — Hybrid Vector + BM25 ​

The goal of retrieval is to "recall the most relevant chunks". Vector-only recall has a known blind spot: queries whose wording overlaps lexically but whose meaning differs — search for "Apple's earnings report" and "apple" may float piles of fruit-related documents to the top. That's why the mature approach is hybrid retrieval with vectors + BM25, fused and re-ordered with RRF (Reciprocal Rank Fusion).

Vector recall:

python
def vector_search(query: str, k: int = 5) -> list[int]:
    """Return the chunk-list indices of the top-k hits."""
    qv = model.encode([query], normalize_embeddings=True)
    scores, ids = index.search(qv.astype("float32"), k)
    return [int(i) for i in ids[0]]

BM25 is the classic sparse retriever built on term frequency / inverse document frequency, and it is razor-sharp on exact keywords (tokenize Chinese text first):

python
import jieba
from rank_bm25 import BM25Okapi

tokenized = [list(jieba.cut(t)) for t in texts]
bm25 = BM25Okapi(tokenized)

def bm25_search(query: str, k: int = 5) -> list[int]:
    scores = bm25.get_scores(list(jieba.cut(query)))
    return scores.argsort()[::-1][:k].tolist()

RRF fusion: ignore the absolute scores and look only at "ranks", merging the ordering information from the two result lists (the hyperparameter 60 is the conventional constant from the RRF paper):

python
def rrf_fuse(rank_lists: list[list[int]], k: int = 60, top_n: int = 5) -> list[int]:
    """Reciprocal Rank Fusion: score each ranked list by 1/(k+rank), then sum across lists."""
    scores: dict[int, float] = {}
    for lst in rank_lists:
        for rank, doc_id in enumerate(lst):
            scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
    return sorted(scores, key=scores.get, reverse=True)[:top_n]

def hybrid_search(query: str, top_n: int = 5) -> list[dict]:
    vec_ids = vector_search(query, 10)     # over-recall in each channel first, then fuse
    bm25_ids = bm25_search(query, 10)
    fused = rrf_fuse([vec_ids, bm25_ids], top_n=top_n)
    return [chunks[i] for i in fused]

Bottom line

Hybrid retrieval almost always beats a single channel. BM25 guards exact keywords, vectors guard semantic relevance, and RRF fusion makes the two complement each other. It's the first free upgrade on the road from "demo" to "usable".

Step 4: Generation — Prompt Assembly ​

The retrieved chunks then get "poured into" the prompt. Assembly follows three rules (see Prompt Engineering for details):

  1. Constrain the knowledge source: answer only from the provided material, and say "I don't know" outright when the material lacks the answer — this is the first gate against hallucination.
  2. Carry the sources: number each chunk and note its source file so every answer is "traceable".
  3. Control the context size: truncate each chunk to a sensible length so you neither blow the context window nor dilute attention with irrelevant text.
python
def build_prompt(query: str, hits: list[dict], max_chars: int = 800) -> str:
    context = "\n\n".join(
        f"[Chunk {i + 1}] (source: {h['source']})\n{h['text'][:max_chars]}"
        for i, h in enumerate(hits)
    )
    return (
        "You are a personal knowledge base assistant. Answer the question using only the material provided below; "
        "if the material does not contain the answer, reply exactly 'The knowledge base contains no relevant information' and do not invent one.\n\n"
        f"### Material\n{context}\n\n### Question\n{query}\n\n### Answer"
    )

def generate(query: str, hits: list[dict]) -> str:
    prompt = build_prompt(query, hits)
    resp = client.chat.completions.create(
        model=LLM_MODEL, messages=[{"role": "user", "content": prompt}],
        temperature=0.3,      # keep the temperature low for knowledge QA to curb improvisation
    )
    return resp.choices[0].message.content

temperature=0.3 is the common setting for knowledge QA — you don't need the model to be "creative", you need it to be "faithful". For more techniques on controlling answer style, see The Prompt Playbook.

Step 5: A Complete, Runnable Demo ​

Wire the first four steps into one self-contained script, rag_demo.py. Usage: drop your personal documents into the kb/ directory, run once with --rebuild to build the index, then just ask:

python
"""
rag_demo.py — personal knowledge base Q&A (minimal pure-Python implementation)

Usage:
    python rag_demo.py --rebuild              # first run: build the index
    python rag_demo.py "What is a vector database?"   # ask a question

Dependencies: pip install sentence-transformers faiss-cpu numpy openai jieba rank-bm25
"""
import json
import sys
from pathlib import Path

import faiss
import jieba
import numpy as np
from openai import OpenAI
from rank_bm25 import BM25Okapi
from sentence_transformers import SentenceTransformer

# ── Configuration ──────────────────────────────────────
KB_DIR = Path("./kb")                      # put your .md / .txt files here
EMBED_MODEL = "BAAI/bge-small-zh-v1.5"     # local embedding model
LLM_MODEL = "gpt-4o-mini"                  # generation model (local Ollama variant at the end)
CHUNK_SIZE, CHUNK_OVERLAP = 500, 50
TOP_N = 4
# ─────────────────────────────────────────────────────────

client = OpenAI()                          # requires the OPENAI_API_KEY environment variable
model = SentenceTransformer(EMBED_MODEL)


def load_texts() -> list[dict]:
    docs = []
    for fp in sorted(KB_DIR.glob("*")):
        if fp.suffix.lower() in {".txt", ".md"}:
            docs.append({"source": str(fp), "text": fp.read_text(encoding="utf-8")})
    return docs


def split_text(text: str, chunk_size=CHUNK_SIZE, overlap=CHUNK_OVERLAP) -> list[str]:
    assert overlap < chunk_size
    if len(text) <= chunk_size:
        return [text]
    chunks, start = [], 0
    while start < len(text):
        end = min(start + chunk_size, len(text))
        if end < len(text):
            for sep in ("\n\n", "\n", "。", ";"):
                pos = text.rfind(sep, start + chunk_size // 2, end)
                if pos != -1:
                    end = pos + len(sep)
                    break
        chunks.append(text[start:end])
        start = end - overlap
    return chunks


def build_index():
    """Chunk → embed → build the FAISS index → persist (metadata + vectors + index)."""
    chunks = [
        {"source": d["source"], "chunk_index": i, "text": c}
        for d in load_texts()
        for i, c in enumerate(split_text(d["text"]))
    ]
    vecs = model.encode([c["text"] for c in chunks], normalize_embeddings=True)
    index = faiss.IndexFlatIP(vecs.shape[1])
    index.add(vecs.astype("float32"))
    faiss.write_index(index, "kb.index")
    np.save("kb_vecs.npy", vecs)
    json.dump(chunks, open("kb_meta.json", "w", encoding="utf-8"), ensure_ascii=False)
    print(f"Indexing complete: {len(chunks)} chunks")


def load_index():
    global chunks, index, bm25
    chunks = json.load(open("kb_meta.json", encoding="utf-8"))
    index = faiss.read_index("kb.index")
    vecs = np.load("kb_vecs.npy")
    bm25 = BM25Okapi([list(jieba.cut(c["text"])) for c in chunks])


def vector_search(q: str, k: int) -> list[int]:
    qv = model.encode([q], normalize_embeddings=True)
    _, ids = index.search(qv.astype("float32"), k)
    return [int(i) for i in ids[0]]


def bm25_search(q: str, k: int) -> list[int]:
    scores = bm25.get_scores(list(jieba.cut(q)))
    return scores.argsort()[::-1][:k].tolist()


def hybrid_search(q: str, top_n: int = TOP_N) -> list[dict]:
    scores: dict[int, float] = {}
    for lst in (vector_search(q, 10), bm25_search(q, 10)):
        for rank, did in enumerate(lst):
            scores[did] = scores.get(did, 0.0) + 1.0 / (60 + rank)
    return [chunks[i] for i in sorted(scores, key=scores.get, reverse=True)[:top_n]]


def ask(query: str) -> str:
    hits = hybrid_search(query)
    context = "\n\n".join(
        f"[Chunk {i + 1}] (source: {h['source']})\n{h['text'][:800]}"
        for i, h in enumerate(hits)
    )
    prompt = (
        "You are a personal knowledge base assistant. Answer the question using only the material below; "
        "if the material contains no answer, reply exactly 'The knowledge base contains no relevant information' and do not invent one.\n\n"
        f"### Material\n{context}\n\n### Question\n{query}\n\n### Answer"
    )
    resp = client.chat.completions.create(
        model=LLM_MODEL, messages=[{"role": "user", "content": prompt}],
        temperature=0.3,
    )
    answer = resp.choices[0].message.content
    # Traceable citations: append the hit chunks to the answer for manual verification
    refs = "\n".join(f"- {h['source']}" for h in hits)
    return f"{answer}\n\n**Sources**:\n{refs}"


if __name__ == "__main__":
    if "--rebuild" in sys.argv:
        build_index()
    load_index()
    query = sys.argv[-1]
    if query in {"--rebuild", "rag_demo.py"}:
        print("Usage: python rag_demo.py \"your question\"")
    else:
        print(ask(query))

Local Ollama variant: swap client and LLM_MODEL for the two lines below, leave every other line untouched, and the whole thing runs fully offline:

python
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
LLM_MODEL = "qwen2.5:7b"     # any local model that supports tool calling; run ollama pull qwen2.5 first

Once it runs you should observe two things: answers carry their sources, and the system can say "the knowledge base contains no relevant information" — exactly the two core improvements RAG offers over a bare LLM.

4. Key Tuning Levers: From "It Runs" to "It Works Well" ​

Almost every RAG problem comes down to "retrieval didn't find the right thing" or "generation didn't use it well". Tune in the order below — the returns diminish as you go:

Tuning leverHow to tuneRule-of-thumb valuesSignal to watch
Chunk sizeExperiment: run one evaluation pass each at 256 / 500 / 800 / 1200Start at 300–800 charactersAnswers missing half the content → too small; answers drifting off topic → too large
OverlapAdjust in step with chunk size10%–20% of the chunk sizeSentences cut off mid-way → increase
top-kIncrease the retrieval count and observe3–8 (recall 10 per channel before fusion)Hits but incomplete answers → increase; context too noisy → decrease
Re-ranking (rerank)Vector recall top 50 → cross-encoder rerank → keep the top 5Always do it (the payoff is large)The right answer keeps landing in the runner-up slot
Source citationsReturn chunk numbers and source filesAttach sources after each answerUsers can't verify the answer → mandatory

Re-ranking is the highest-value step. Vector retrieval is "approximate matching"; a cross-encoder reranker (such as BAAI/bge-reranker-v2-m3) scores each "query + candidate chunk" pair through the model as a whole — far more accurate but slower, so "coarse recall first, precise rerank second" is the standard play:

python
from sentence_transformers import CrossEncoder

reranker = CrossEncoder("BAAI/bge-reranker-v2-m3")
cands = vector_search(query, 50)                       # coarse recall first
pairs = [(query, chunks[i]["text"]) for i in cands]
scores = reranker.predict(pairs)
top = [cands[i] for i in scores.argsort()[::-1][:TOP_N]]  # then precise re-ranking

5. Evaluation: Don't Trust "It Feels Good" ​

Evaluate RAG in two segments: the retrieval segment (did it find the right things?) and the generation segment (did it answer correctly?). Without evaluation you will never know whether a chunk tweak made things better or worse — this is the dividing line between dabbling and engineering.

Retrieval metrics (they require labeling "which chunks are correct for each question" by hand, or generating those labels automatically from document structure):

MetricFormulaMeaning
Recall@kcorrect chunks hit / total correct chunksHow many of the correct chunks were recovered among the top k
MRRReciprocal of the first correct result's rank, averagedHow smoothly "the first hit is the right one"
Hit@kWhether at least one correct chunk appears in the top kWhether it's good enough to work with

Generation metrics: faithfulness (can the answer be supported by the retrieved chunks?) and answer relevance (does it address the actual question?). The usual approach is LLM-as-judge, letting a second model do the scoring — see Building an LLM Evaluation Suite for a concrete build guide, and LLM Evaluation and Benchmarks for the full theory behind the metric system.

Evaluation discipline

  • Build the eval set before you tune: collect 30–100 items of "question + correct chunks + reference answer". Every tuning decision made without an eval set is self-consolation.
  • Measure retrieval and generation separately: however good generation is, it's wasted if retrieval recall is zero; optimize the retrieval metrics on their own first, then move on to generation.
  • Change one variable, re-run: tweak chunk size, top-k, and rerank all at once and a good result tells you nothing about which change earned it.

6. Common Pitfalls ​

PitfallSymptomRoot causeFix
Chunking breaks semanticsAnswers missing half the content, or contradicting themselvesThe fixed window lands right in the middle of a paragraph or sentenceAdd overlap, fall back to natural boundaries, split on paragraphs
Retrieval misses → hallucinationThe model fabricates with a straight faceRelevant chunks were never recalled, so the model has to wing itHybrid retrieval, larger recall, rerank, and a prompt that explicitly orders "say so when you don't know"
Unfiltered boilerplateRetrieval keeps hitting footers / copyright notices / tables of contentsFooter text gets mixed into chunks and indexedStrip boilerplate during cleaning and add filtering rules
Embeddings not normalizedSimilarity scores aren't comparableCosine similarity and inner product used interchangeablyStandardize on normalize_embeddings=True
Context truncatedAnswers stop halfway, or later material gets ignoredChunk count × chunk length exceeds the windowTruncate individual chunks, control top-k, or move to a long-context model as needed
Index never refreshedNew documents can't be foundDocuments were edited but the index was never rebuiltFold index building into the document update workflow (CI or a scheduled job)
Testing only one or two questions"Feels" great, falls apart in productionToo small a sample that all happens to hitBuild a 30+ item eval set and track quantitative metrics
Retrieval hits but the answer is wrongThe material is right there, yet the answer is wrongAn unconstrained prompt, or distractors mixed into the chunksCheck the prompt constraints, check chunk purity, tune rerank

For more general engineering anti-patterns, see Common Pitfalls and Anti-Patterns.

7. Where to Go Next ​

Once the basic RAG runs and has been evaluated, level up as needed:

  1. GraphRAG: multi-hop and global questions. Basic RAG is weak on "relational questions spanning multiple chunks" (such as "which documents all mention the same event"). Extracting the documents into an entity–relationship knowledge graph and retrieving over it supports multi-hop reasoning — see Knowledge Graphs and Knowledge Injection for the underlying principles.
  2. Agentic RAG: turn retrieval into a tool. Let the agent decide on its own "when to retrieve, how many times, whether to ask a clarifying question, whether to rewrite the query". Multi-turn clarification markedly improves vague questions, and query rewriting eases the mismatch between "how users ask" and "how documents are written" — this is the tool-calling loop from Building an Agent from Scratch applied to RAG; see Agent Core Concepts for the conceptual foundation.
  3. Streaming and caching. Stream generation output over SSE for a better experience; add semantic caching for high-frequency repeated questions (similar questions hit the cached answer directly), saving both tokens and latency — see Inference Optimization and Quantization and Deployment and Inference Optimization in Practice.
  4. Embedding fine-tuning. When domain jargon is dense and general-purpose embeddings can't tell things apart, fine-tune the embedding model on domain data — see Fine-Tuning and PEFT.

Closing verdict

A "passing" RAG system: stable metrics on an eval set, answers that carry citations, and an honest "I don't know" when the material doesn't cover it. Nail those three and it already outperforms most fake RAG setups of "bare LLM + a folder of knowledge files". Look up any unfamiliar term in the Glossary; for the big picture, see What Are the Hot AI Concepts.

Further Reading ​

References ​