Skip to content

Vector Databases and Semantic Search

At a glance Vector databases store embeddings and retrieve by semantic similarity via approximate nearest neighbor search, forming the storage backbone of RAG systems; this page covers embedding basics, semantic vs. keyword search, similarity metrics, ANN indexes, product selection, and engineering pitfalls.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Vector Databases and Semantic Search ​

A vector database is a database purpose-built for storing and retrieving high-dimensional vectors (embeddings). It supports approximate nearest neighbor (ANN) search by semantic similarity, and it is the storage backbone of every RAG system.

The core problem it solves is one traditional databases can't: what goes in isn't a structured record like "who is John," but a high-dimensional coordinate capturing "what this text / image / audio clip means semantically." And retrieval doesn't rely on exact WHERE matching — you ask "which 10 items are most similar to this query?" Who uses it? Nearly every production RAG application, from AI search engines like Perplexity to internal enterprise document Q&A, has a vector database underneath; recommendation systems, multimodal search, and agent memory are all its home turf. Before you understand the database, understand what it stores: the embedding.

Where it fits

If RAG is the architecture that "bolts an external knowledge base onto an LLM," then the vector database is the storage and retrieval engine of that knowledge base. Whether knowledge can be found at all — and how accurately — depends largely on this layer.

1. Background: Embeddings and Vectors ​

An embedding maps unstructured data — text, images, audio — into a high-dimensional vector. The result is a vector of a few hundred to a few thousand floats, e.g. [0.13, -0.42, 0.87, ...]. Its defining property is a geometric assumption: objects with similar semantics sit close together in vector space.

  • Text embeddings: a whole sentence or paragraph encoded into one vector, produced by the encoder of models like Transformer. Common dimensions: 768 (BERT-base), 1024 (BGE), 1536 (OpenAI text-embedding-3-small), 3072 (text-embedding-3-large).
  • Image embeddings: images encoded through the vision branch of models like CLIP, able to share a space with text vectors (see Multimodal Models).
  • Each vector is just a point in high-dimensional space; the closer two points are (under some distance metric), the more semantically similar they are.
ConceptOne-line explanationTypical dimensions
Word vectors (Word2Vec)One vector per word; "king − man + woman ≈ queen"100–300
Sentence/text vectorsOne vector per sentence/paragraph, semantically comparable768–3072
Multimodal vectorsText and images mapped into one shared vector space512–1024

The one-line verdict: embedding quality caps retrieval quality. A vector database just faithfully runs nearest-neighbor search; if the embedding model itself can't separate meanings (say, "Apple the company" vs. "apple the fruit"), no indexing trick will save you.

Keyword retrieval in traditional search engines is typified by BM25: split documents into terms, weigh term frequency against inverse document frequency, and score query-document pairs by weighted shared terms. It has no idea what the words mean.

DimensionBM25 keyword searchVector semantic search
Matching unitLiteral terms (tokens)Semantic vectors
Understands synonymsNo ("car" won't match "automobile")Yes (semantic proximity is enough)
Multilingual abilityWeak (tokenization issues)Strong (compare directly in one vector space)
Query robustnessRephrase and you find nothingDifferent wording, same meaning, still recalled
PrecisionHigh (a hit is relevant)Medium (may recall "similar but irrelevant")
Needs an inverted indexYesNo (needs an ANN index)
Typical latencyMillisecondsMilliseconds (approximate)
Best forExact keywords, IDs, logsNatural-language Q&A, fuzzy semantics

A quick example: a user asks "how do I water my flowers." BM25 only hits if a document literally contains "water my flowers"; vector search can recall an article on "plant moisture management" — because the two vectors sit close together.

The one-line verdict: don't worship vector search as all-powerful. For order numbers, ID numbers, and exact product codes, BM25 wins outright; for open-domain Q&A and fuzzy semantics, vectors win outright. That's why mature systems almost always use hybrid retrieval (see Section 6) to stack the two, rather than picking one.

3. Similarity Metrics: Cosine Similarity, Dot Product, Euclidean Distance ​

Given a query vector q and a candidate vector d, three metrics dominate:

MetricFormulaRangeWhen to use
Cosine similarity$\cos\theta=\frac{q\cdot d}{|q||d|}$[-1, 1] (higher = more similar)The most common default; cares about "direction" rather than magnitude, the most robust choice for semantic comparison of embeddings
Dot product (inner product)$q\cdot d=\sum_i q_i d_i$UnboundedWhen vectors are already normalized or the model was trained with inner product; fits explicitly modeled "relevance scores"
Euclidean distance (L2)$\sqrt{\sum_i (q_i-d_i)^2}$[0, +∞) (smaller = closer)When you genuinely need "spatial distance" semantics; smaller L2 = more similar, sort ascending

Metric–index mismatches

Picking the wrong metric silently degrades recall. Classic cases: the index was built with "cosine" but the query sends "dot product" scores; or you use L2 distance but forget the embeddings aren't normalized, systematically favoring large-magnitude vectors. Most text embedding models recommend L2 normalization first, then dot product/cosine — with normalized vectors all three metrics produce identical rankings, and switching between them becomes trivial.

Rules of thumb: with no special reason, use cosine similarity (text scenarios); image/video features often use dot product; reach for Euclidean distance only when you truly need "spatial distance" semantics. Whichever you pick, keep it consistent across the whole pipeline and validate on an evaluation set (see LLM Evaluation and Benchmarks).

4. The Core Technology: ANN Indexes ​

The naive approach is brute force: compute similarity between the query vector and every vector in the database, then sort. At ten-thousand scale that's fine; at millions to billions of vectors, a single query means millions of dot products — unacceptable latency. Hence Approximate Nearest Neighbor (ANN) indexes: trade a sliver of recall for orders-of-magnitude speed.

ApproachSpeedRecallMemoryScale
Exact (brute force)Slow (O(N))100%Low< 100K
ANN (HNSW/IVF/PQ)Very fast (milliseconds)90–99%Medium–highMillions to billions

The three mainstream ANN algorithms:

AlgorithmFull name / ideaProsCons
HNSWHierarchical Navigable Small World graphHigh recall, fast queries, no training neededMemory-hungry, costly inserts
IVFInverted File: cluster into buckets first, then search within bucketsControllable memory, fast to buildRecall loss at bucket boundaries, parameter-sensitive
PQProduct QuantizationCompresses vectors massively (8–32× memory savings)Precision loss, table-lookup queries are complex
text
HNSW search process (schematic):

 Layer 4  ──●────────────────────  entry point (sparsest, long hops)
          │
 Layer 3  ──●──●──●──────────────
          │        │
 Layer 2  ──●──●──●──●──●────────  greedy search, descending layer by layer
          │     │     │
 Layer 1  ──●──●──●──●──●──●──●─  bottom layer (densest)

 The query enters at the top-layer entry point → greedily moves to the nearest
 neighbor at each layer → at the bottom layer, beam search collects the Top-K candidates

What happens in the diagram: HNSW organizes vectors into a multi-layer graph — upper layers are sparse and handle "long-distance jumps," the bottom layer is dense and handles "precise localization." The query starts at the top-layer entry point, moves toward "the currently nearest neighbor" at each layer, and descends to the bottom, where it collects and ranks a batch of candidates. The whole process is greedy, so in theory it can miss the true global nearest point — that's where the "approximate" comes from; in practice recall loss is usually < 1%.

  • HNSW is today's default workhorse (Milvus, Qdrant, pgvector, and FAISS all support it); Rust/Go implementations typically deliver p95 query latency from a few milliseconds to a few tens of milliseconds;
  • IVF + PQ combinations are common at billion scale when memory is tight: quantize and compress first, then search within buckets;
  • The one-line verdict: default to HNSW; consider IVF/PQ once you hit hundreds of millions of vectors or hit memory limits. The original paper is in the References (HNSW, arXiv:1603.09320).

5. Products and Selection ​

As of mid-2025, the vector search ecosystem splits into two camps: vector databases (complete database products) and vector search libraries/extensions (embedded into applications or existing databases). The main options:

ProductTypeKey featuresANN supportDeployment
FAISS (Meta, open source)LibraryThe most mature low-level algorithm library, strong GPU acceleration, the de facto standardHNSW/IVF/PQEmbedded in your app (manage the data yourself)
Milvus (Zilliz)Distributed vector databaseCloud-native, billion-scale, GPU index (CAGRA), strong filteringHNSW/IVF/PQ etc.Local/self-hosted/cloud
Qdrant (open source)Vector databaseRust implementation, excellent filtering and payloads, low latencyHNSWLocal/self-hosted/cloud
PineconeManaged vector databaseFully managed, zero ops, ServerlessProprietary optimized indexCloud only
WeaviateVector databaseBuilt-in GraphQL, generative modules, modular designHNSWLocal/cloud
pgvectorPostgreSQL extensionCoexists with existing SQL, transactional consistency, zero new infrastructureHNSW/IVFShips with PostgreSQL
ChromaLightweight vector databasePython-native, fastest to get startedHNSWLocal/embedded
ElasticsearchSearch engine + vectorsOne system for BM25/inverted search plus vectors, via the dense_vector fieldHNSWLocal/cloud

Self-hosted vs. managed:

DimensionLocal / self-hostedManaged (Pinecone, cloud Milvus/Qdrant)
Initial costLow (open source, free)High (pay as you go)
Ops burdenHigh (index tuning, scaling, backups)Low (the platform handles it)
Data complianceData stays in-houseReview data-residency and privacy terms
Scale elasticityPlan it yourselfElastic scale up/down
Best forLearning, POCs, on-prem deployments, teams with K8s expertiseFast launch, spiky workloads, small teams

The one-line verdict: if your team has no dedicated search engineer, start with a managed service and move back to self-hosting when scale and cost pressure arrive. Already running PostgreSQL with under a few million vectors? Try pgvector first — zero new components, transactions and SQL for free; it's the cheapest on-ramp.

6. The Role in the RAG Pipeline ​

In a typical RAG pipeline (see Build a RAG App from Scratch), the vector database handles storage and recall:

text
Document chunks → (1) embedding → (2) write to vector database (index)
                                        │
User question → (3) embedding → (4) ANN retrieval Top-K → (5) rerank
                                        │
        Metadata filtering (time/source/permissions) applies at (4)
                                        ▼
        Retrieved results assembled into the prompt → LLM generates the answer

Three key practices:

  1. Index → retrieve → rerank: the vector database only recalls a "wide" candidate set (usually Top-50–200); a reranker (e.g. cross-encoder) then narrows it to Top-5–10 for the LLM. Recall wide enough to miss nothing; rerank tightly enough to feed only the best — the highest-leverage lever for answer quality;
  2. Metadata filtering: nearly every vector database supports stacking filters on queries (time range, document source, business line, permission tags). Filter first, then run ANN search: a far smaller candidate set and better accuracy. Permission filtering matters most — otherwise documents a user "can find but can't read" leak into results;
  3. Hybrid retrieval (dense vectors + sparse BM25): merge vector and BM25 results with weighted fusion (e.g. RRF, reciprocal rank fusion) to cover both semantics and exact matching. It's the standard play behind AI search products like Perplexity, and it significantly mitigates vector search's "similar but irrelevant" failure mode.

Evaluate first

Before pairing a vector database with RAG, establish a retrieval quality baseline: prepare a labeled set of "question → expected document" pairs and measure recall@k (was the wanted document recalled?) and hit rate@k (were the recalled documents relevant?). Without this baseline, tuning index parameters, embedding models, and hybrid weights is guesswork.

7. Advanced Topics ​

7.1 Vectors + Knowledge Graphs: Complementary, Not Competing ​

Vectors excel at "fuzzy semantic association" but know nothing about "relations and rules"; knowledge graphs excel at exact multi-hop relational reasoning but are fragile against natural-language variation. They complement each other (see Knowledge Graphs and Knowledge Injection):

ScenarioVector databaseKnowledge graph
"Which documents discuss microservice degradation?"Strong (semantic recall)Weak
"Which services do the services that service A depends on depend on?"WeakStrong (multi-hop reasoning)
Factual consistencyNot guaranteed (vectors are probabilistic)Verifiable
Permissions/rulesBuild the filtering yourselfNaturally rule-friendly

The pattern: vector search recalls candidate entities first, the graph then validates relations and reasons over paths, and finally the answer is assembled.

7.2 Multimodal Vectors: Cross-Modal Retrieval ​

CLIP-style models map text and images into the same vector space, so "upload a photo, find similar products" and "describe something in words, find images/video clips" all reduce to a single nearest-neighbor query (see Multimodal Models). The vector database imposes only one hard requirement on multimodal search: all embeddings must come from the same model family — mixing vectors from different models is like comparing distances across different coordinate systems; the results are meaningless.

7.3 Vector Databases as Agent Memory ​

An AI agent has a limited context window, so it needs to persist past interactions, tool outputs, and user preferences into "queryable memory." The common recipe: encode each interaction as a vector and store it in a vector database; when a new task arrives, recall relevant past experience by similarity and stuff it into the context (see Build an Agent from Scratch). This is technically almost isomorphic to RAG; the difference is that you're writing process records, not knowledge documents, so the bar is higher for write frequency, deduplication, and freshness (memory expiration).

8. Engineering Pitfalls ​

The Curse of Dimensionality ​

The higher the dimension, the more the distances between arbitrary points converge, and nearest-neighbor discrimination degrades. Mitigations: pick embedding dimensions matched to your data volume (384 dimensions is plenty for small datasets), reduce dimensions with PCA/quantization, and prefer cosine over Euclidean distance.

Cold Start ​

Fresh data has too few neighbors, so retrieval quality suffers; or an embedding model upgrade makes old and new vector spaces inconsistent. Mitigations: allow hybrid retrieval as a safety net during cold start; when upgrading embedding models, re-embed everything and rebuild the index — never mix old and new vectors.

Updates and Deletes ​

An "update" in a vector index usually means delete old + insert new, not in-place modification; frequent writes and deletes slow HNSW down (the graph structure must be maintained). In distributed implementations, deletes are usually "soft delete + periodic compaction." Mitigations: batch your writes; assess throughput needs before choosing an architecture (the read/write ratio determines index configuration).

Cost ​

The money goes to memory (the HNSW graph lives entirely in RAM) and embedding compute. One million 1536-dimension vectors take roughly 6 GB of memory (excluding index overhead). Mitigations: compress memory with PQ quantization, tier hot and cold data, cache per-turn query embeddings, and apply the same cost mindset as in Inference Optimization and Quantization.

A vector database is not a silver bullet

It solves "find similar by semantics," not "is the answer correct and factually consistent." Dumping all company data into a vector store and expecting the LLM to cite everything correctly conflates retrieval recall with fact verification — the latter takes evaluation, reranking, and guardrails. More failure stories in Common Pitfalls and Anti-Patterns.

Further Reading ​

References ​