Appearance
Vector Databases and Semantic Search
A vector database is a database purpose-built for storing and retrieving high-dimensional vectors (embeddings). It supports approximate nearest neighbor (ANN) search by semantic similarity, and it is the storage backbone of every RAG system.
The core problem it solves is one traditional databases can't: what goes in isn't a structured record like "who is John," but a high-dimensional coordinate capturing "what this text / image / audio clip means semantically." And retrieval doesn't rely on exact WHERE matching — you ask "which 10 items are most similar to this query?" Who uses it? Nearly every production RAG application, from AI search engines like Perplexity to internal enterprise document Q&A, has a vector database underneath; recommendation systems, multimodal search, and agent memory are all its home turf. Before you understand the database, understand what it stores: the embedding.
Where it fits
If RAG is the architecture that "bolts an external knowledge base onto an LLM," then the vector database is the storage and retrieval engine of that knowledge base. Whether knowledge can be found at all — and how accurately — depends largely on this layer.
1. Background: Embeddings and Vectors
An embedding maps unstructured data — text, images, audio — into a high-dimensional vector. The result is a vector of a few hundred to a few thousand floats, e.g. [0.13, -0.42, 0.87, ...]. Its defining property is a geometric assumption: objects with similar semantics sit close together in vector space.
- Text embeddings: a whole sentence or paragraph encoded into one vector, produced by the encoder of models like Transformer. Common dimensions: 768 (BERT-base), 1024 (BGE), 1536 (OpenAI text-embedding-3-small), 3072 (text-embedding-3-large).
- Image embeddings: images encoded through the vision branch of models like CLIP, able to share a space with text vectors (see Multimodal Models).
- Each vector is just a point in high-dimensional space; the closer two points are (under some distance metric), the more semantically similar they are.
| Concept | One-line explanation | Typical dimensions |
|---|---|---|
| Word vectors (Word2Vec) | One vector per word; "king − man + woman ≈ queen" | 100–300 |
| Sentence/text vectors | One vector per sentence/paragraph, semantically comparable | 768–3072 |
| Multimodal vectors | Text and images mapped into one shared vector space | 512–1024 |
The one-line verdict: embedding quality caps retrieval quality. A vector database just faithfully runs nearest-neighbor search; if the embedding model itself can't separate meanings (say, "Apple the company" vs. "apple the fruit"), no indexing trick will save you.
2. Semantic Search vs. Keyword Search
Keyword retrieval in traditional search engines is typified by BM25: split documents into terms, weigh term frequency against inverse document frequency, and score query-document pairs by weighted shared terms. It has no idea what the words mean.
| Dimension | BM25 keyword search | Vector semantic search |
|---|---|---|
| Matching unit | Literal terms (tokens) | Semantic vectors |
| Understands synonyms | No ("car" won't match "automobile") | Yes (semantic proximity is enough) |
| Multilingual ability | Weak (tokenization issues) | Strong (compare directly in one vector space) |
| Query robustness | Rephrase and you find nothing | Different wording, same meaning, still recalled |
| Precision | High (a hit is relevant) | Medium (may recall "similar but irrelevant") |
| Needs an inverted index | Yes | No (needs an ANN index) |
| Typical latency | Milliseconds | Milliseconds (approximate) |
| Best for | Exact keywords, IDs, logs | Natural-language Q&A, fuzzy semantics |
A quick example: a user asks "how do I water my flowers." BM25 only hits if a document literally contains "water my flowers"; vector search can recall an article on "plant moisture management" — because the two vectors sit close together.
The one-line verdict: don't worship vector search as all-powerful. For order numbers, ID numbers, and exact product codes, BM25 wins outright; for open-domain Q&A and fuzzy semantics, vectors win outright. That's why mature systems almost always use hybrid retrieval (see Section 6) to stack the two, rather than picking one.
3. Similarity Metrics: Cosine Similarity, Dot Product, Euclidean Distance
Given a query vector q and a candidate vector d, three metrics dominate:
| Metric | Formula | Range | When to use |
|---|---|---|---|
| Cosine similarity | $\cos\theta=\frac{q\cdot d}{|q||d|}$ | [-1, 1] (higher = more similar) | The most common default; cares about "direction" rather than magnitude, the most robust choice for semantic comparison of embeddings |
| Dot product (inner product) | $q\cdot d=\sum_i q_i d_i$ | Unbounded | When vectors are already normalized or the model was trained with inner product; fits explicitly modeled "relevance scores" |
| Euclidean distance (L2) | $\sqrt{\sum_i (q_i-d_i)^2}$ | [0, +∞) (smaller = closer) | When you genuinely need "spatial distance" semantics; smaller L2 = more similar, sort ascending |
Metric–index mismatches
Picking the wrong metric silently degrades recall. Classic cases: the index was built with "cosine" but the query sends "dot product" scores; or you use L2 distance but forget the embeddings aren't normalized, systematically favoring large-magnitude vectors. Most text embedding models recommend L2 normalization first, then dot product/cosine — with normalized vectors all three metrics produce identical rankings, and switching between them becomes trivial.
Rules of thumb: with no special reason, use cosine similarity (text scenarios); image/video features often use dot product; reach for Euclidean distance only when you truly need "spatial distance" semantics. Whichever you pick, keep it consistent across the whole pipeline and validate on an evaluation set (see LLM Evaluation and Benchmarks).
4. The Core Technology: ANN Indexes
The naive approach is brute force: compute similarity between the query vector and every vector in the database, then sort. At ten-thousand scale that's fine; at millions to billions of vectors, a single query means millions of dot products — unacceptable latency. Hence Approximate Nearest Neighbor (ANN) indexes: trade a sliver of recall for orders-of-magnitude speed.
| Approach | Speed | Recall | Memory | Scale |
|---|---|---|---|---|
| Exact (brute force) | Slow (O(N)) | 100% | Low | < 100K |
| ANN (HNSW/IVF/PQ) | Very fast (milliseconds) | 90–99% | Medium–high | Millions to billions |
The three mainstream ANN algorithms:
| Algorithm | Full name / idea | Pros | Cons |
|---|---|---|---|
| HNSW | Hierarchical Navigable Small World graph | High recall, fast queries, no training needed | Memory-hungry, costly inserts |
| IVF | Inverted File: cluster into buckets first, then search within buckets | Controllable memory, fast to build | Recall loss at bucket boundaries, parameter-sensitive |
| PQ | Product Quantization | Compresses vectors massively (8–32× memory savings) | Precision loss, table-lookup queries are complex |
text
HNSW search process (schematic):
Layer 4 ──●──────────────────── entry point (sparsest, long hops)
│
Layer 3 ──●──●──●──────────────
│ │
Layer 2 ──●──●──●──●──●──────── greedy search, descending layer by layer
│ │ │
Layer 1 ──●──●──●──●──●──●──●─ bottom layer (densest)
The query enters at the top-layer entry point → greedily moves to the nearest
neighbor at each layer → at the bottom layer, beam search collects the Top-K candidatesWhat happens in the diagram: HNSW organizes vectors into a multi-layer graph — upper layers are sparse and handle "long-distance jumps," the bottom layer is dense and handles "precise localization." The query starts at the top-layer entry point, moves toward "the currently nearest neighbor" at each layer, and descends to the bottom, where it collects and ranks a batch of candidates. The whole process is greedy, so in theory it can miss the true global nearest point — that's where the "approximate" comes from; in practice recall loss is usually < 1%.
- HNSW is today's default workhorse (Milvus, Qdrant, pgvector, and FAISS all support it); Rust/Go implementations typically deliver p95 query latency from a few milliseconds to a few tens of milliseconds;
- IVF + PQ combinations are common at billion scale when memory is tight: quantize and compress first, then search within buckets;
- The one-line verdict: default to HNSW; consider IVF/PQ once you hit hundreds of millions of vectors or hit memory limits. The original paper is in the References (HNSW, arXiv:1603.09320).
5. Products and Selection
As of mid-2025, the vector search ecosystem splits into two camps: vector databases (complete database products) and vector search libraries/extensions (embedded into applications or existing databases). The main options:
| Product | Type | Key features | ANN support | Deployment |
|---|---|---|---|---|
| FAISS (Meta, open source) | Library | The most mature low-level algorithm library, strong GPU acceleration, the de facto standard | HNSW/IVF/PQ | Embedded in your app (manage the data yourself) |
| Milvus (Zilliz) | Distributed vector database | Cloud-native, billion-scale, GPU index (CAGRA), strong filtering | HNSW/IVF/PQ etc. | Local/self-hosted/cloud |
| Qdrant (open source) | Vector database | Rust implementation, excellent filtering and payloads, low latency | HNSW | Local/self-hosted/cloud |
| Pinecone | Managed vector database | Fully managed, zero ops, Serverless | Proprietary optimized index | Cloud only |
| Weaviate | Vector database | Built-in GraphQL, generative modules, modular design | HNSW | Local/cloud |
| pgvector | PostgreSQL extension | Coexists with existing SQL, transactional consistency, zero new infrastructure | HNSW/IVF | Ships with PostgreSQL |
| Chroma | Lightweight vector database | Python-native, fastest to get started | HNSW | Local/embedded |
| Elasticsearch | Search engine + vectors | One system for BM25/inverted search plus vectors, via the dense_vector field | HNSW | Local/cloud |
Self-hosted vs. managed:
| Dimension | Local / self-hosted | Managed (Pinecone, cloud Milvus/Qdrant) |
|---|---|---|
| Initial cost | Low (open source, free) | High (pay as you go) |
| Ops burden | High (index tuning, scaling, backups) | Low (the platform handles it) |
| Data compliance | Data stays in-house | Review data-residency and privacy terms |
| Scale elasticity | Plan it yourself | Elastic scale up/down |
| Best for | Learning, POCs, on-prem deployments, teams with K8s expertise | Fast launch, spiky workloads, small teams |
The one-line verdict: if your team has no dedicated search engineer, start with a managed service and move back to self-hosting when scale and cost pressure arrive. Already running PostgreSQL with under a few million vectors? Try pgvector first — zero new components, transactions and SQL for free; it's the cheapest on-ramp.
6. The Role in the RAG Pipeline
In a typical RAG pipeline (see Build a RAG App from Scratch), the vector database handles storage and recall:
text
Document chunks → (1) embedding → (2) write to vector database (index)
│
User question → (3) embedding → (4) ANN retrieval Top-K → (5) rerank
│
Metadata filtering (time/source/permissions) applies at (4)
▼
Retrieved results assembled into the prompt → LLM generates the answerThree key practices:
- Index → retrieve → rerank: the vector database only recalls a "wide" candidate set (usually Top-50–200); a reranker (e.g. cross-encoder) then narrows it to Top-5–10 for the LLM. Recall wide enough to miss nothing; rerank tightly enough to feed only the best — the highest-leverage lever for answer quality;
- Metadata filtering: nearly every vector database supports stacking filters on queries (time range, document source, business line, permission tags). Filter first, then run ANN search: a far smaller candidate set and better accuracy. Permission filtering matters most — otherwise documents a user "can find but can't read" leak into results;
- Hybrid retrieval (dense vectors + sparse BM25): merge vector and BM25 results with weighted fusion (e.g. RRF, reciprocal rank fusion) to cover both semantics and exact matching. It's the standard play behind AI search products like Perplexity, and it significantly mitigates vector search's "similar but irrelevant" failure mode.
Evaluate first
Before pairing a vector database with RAG, establish a retrieval quality baseline: prepare a labeled set of "question → expected document" pairs and measure recall@k (was the wanted document recalled?) and hit rate@k (were the recalled documents relevant?). Without this baseline, tuning index parameters, embedding models, and hybrid weights is guesswork.
7. Advanced Topics
7.1 Vectors + Knowledge Graphs: Complementary, Not Competing
Vectors excel at "fuzzy semantic association" but know nothing about "relations and rules"; knowledge graphs excel at exact multi-hop relational reasoning but are fragile against natural-language variation. They complement each other (see Knowledge Graphs and Knowledge Injection):
| Scenario | Vector database | Knowledge graph |
|---|---|---|
| "Which documents discuss microservice degradation?" | Strong (semantic recall) | Weak |
| "Which services do the services that service A depends on depend on?" | Weak | Strong (multi-hop reasoning) |
| Factual consistency | Not guaranteed (vectors are probabilistic) | Verifiable |
| Permissions/rules | Build the filtering yourself | Naturally rule-friendly |
The pattern: vector search recalls candidate entities first, the graph then validates relations and reasons over paths, and finally the answer is assembled.
7.2 Multimodal Vectors: Cross-Modal Retrieval
CLIP-style models map text and images into the same vector space, so "upload a photo, find similar products" and "describe something in words, find images/video clips" all reduce to a single nearest-neighbor query (see Multimodal Models). The vector database imposes only one hard requirement on multimodal search: all embeddings must come from the same model family — mixing vectors from different models is like comparing distances across different coordinate systems; the results are meaningless.
7.3 Vector Databases as Agent Memory
An AI agent has a limited context window, so it needs to persist past interactions, tool outputs, and user preferences into "queryable memory." The common recipe: encode each interaction as a vector and store it in a vector database; when a new task arrives, recall relevant past experience by similarity and stuff it into the context (see Build an Agent from Scratch). This is technically almost isomorphic to RAG; the difference is that you're writing process records, not knowledge documents, so the bar is higher for write frequency, deduplication, and freshness (memory expiration).
8. Engineering Pitfalls
The Curse of Dimensionality
The higher the dimension, the more the distances between arbitrary points converge, and nearest-neighbor discrimination degrades. Mitigations: pick embedding dimensions matched to your data volume (384 dimensions is plenty for small datasets), reduce dimensions with PCA/quantization, and prefer cosine over Euclidean distance.
Cold Start
Fresh data has too few neighbors, so retrieval quality suffers; or an embedding model upgrade makes old and new vector spaces inconsistent. Mitigations: allow hybrid retrieval as a safety net during cold start; when upgrading embedding models, re-embed everything and rebuild the index — never mix old and new vectors.
Updates and Deletes
An "update" in a vector index usually means delete old + insert new, not in-place modification; frequent writes and deletes slow HNSW down (the graph structure must be maintained). In distributed implementations, deletes are usually "soft delete + periodic compaction." Mitigations: batch your writes; assess throughput needs before choosing an architecture (the read/write ratio determines index configuration).
Cost
The money goes to memory (the HNSW graph lives entirely in RAM) and embedding compute. One million 1536-dimension vectors take roughly 6 GB of memory (excluding index overhead). Mitigations: compress memory with PQ quantization, tier hot and cold data, cache per-turn query embeddings, and apply the same cost mindset as in Inference Optimization and Quantization.
A vector database is not a silver bullet
It solves "find similar by semantics," not "is the answer correct and factually consistent." Dumping all company data into a vector store and expecting the LLM to cite everything correctly conflates retrieval recall with fact verification — the latter takes evaluation, reranking, and guardrails. More failure stories in Common Pitfalls and Anti-Patterns.
Further Reading
- Retrieval-Augmented Generation (RAG) — the core application of vector databases
- Build a RAG App from Scratch — end-to-end indexing, retrieval, and reranking
- Knowledge Graphs and Knowledge Injection — pairing vectors with graphs
- Multimodal Models — the foundation of cross-modal vector retrieval
- AI Agents — vector databases as long-term memory
- Perplexity and AI Search — hybrid retrieval in a flagship product
- Inference Optimization and Quantization — cutting costs on the model side
- Common Pitfalls and Anti-Patterns — retrieval and RAG engineering traps
- Glossary — quick reference for vector, ANN, and related terms
- Curated Resources — vector database and retrieval tooling
References
- Malkov & Yashunin. Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs (the HNSW paper) — the most influential work in ANN and the original source of HNSW
- Johnson, Douze & Jégou. Billion-scale similarity search with GPUs (the FAISS paper) — the official paper behind the FAISS library
- FAISS official documentation (GitHub) — Meta's open-source vector search library
- Milvus official documentation — the open-source distributed vector database
- Qdrant official documentation — the Rust-implemented vector database
- pgvector (GitHub) — the PostgreSQL vector extension
- Chroma official documentation — a Python-native lightweight vector database
- Pinecone Learning Center: vector database fundamentals — a beginner-friendly vector retrieval tutorial