Every large language model has an expiration date baked in. Its knowledge stops at the end of training: it cannot tell you what happened yesterday, what is in your company's private wiki, or what your product's refund policy says. Retrieval-Augmented Generation (RAG) fixes this by giving the model a memory it can look things up in — and the vector database is the machinery that makes the looking-up fast. This article explains embeddings, similarity search, approximate indexes, and the full RAG pipeline, with the intuition behind each piece.
The problem RAG solves
A language model stores knowledge the hard way: smeared across billions of weights during training. That gives you three problems. First, the knowledge cutoff — anything after training is unknown. Second, hallucination — when the model does not know something, it generates the most plausible-sounding answer anyway. Third, private data — your internal documents were never in the training set, and retraining a model every time a document changes is absurd.
RAG sidesteps all three with a simple idea, introduced by Lewis and colleagues at Meta in 2020: instead of forcing the model to memorize facts, retrieve the relevant documents at query time and hand them to the model as context. The model still does the reasoning and language generation, but the facts come from your data, fresh and attributable. The hard part is the retrieval: given a question, find the few relevant passages among millions, in milliseconds. That is what vector databases are built for.
Embeddings: meaning as coordinates
A vector database does not store text the way a filing cabinet stores paper. It stores embeddings: dense numerical vectors, typically a few hundred to a few thousand dimensions, produced by a neural network trained so that texts with similar meaning land near each other in vector space.
The intuition is geometric. Take three sentences:
- “The cat sat on the mat.”
- “A feline rested on the rug.”
- “Quantum computers use qubits.”
An embedding model maps the first two to nearby points — different words, same meaning — and the third to a distant point. A traditional database sees three unrelated strings; an embedding sees that two of them are neighbours. Common embedding models produce 384 dimensions (lightweight models like MiniLM), 768, or 1536 (OpenAI's text-embedding-3-small). Each dimension captures some learned aspect of meaning, though individual dimensions are not human-interpretable — only the geometry of the whole space matters.
This is the foundational trick of the entire system: semantic similarity becomes spatial proximity. “Find me documents about X” becomes “find me vectors near the vector for X.”
Finding neighbours: cosine similarity
Given a query vector q and a document vector d, how do we measure “nearness”? The workhorse is cosine similarity: the cosine of the angle between the two vectors.
cos(q, d) = (q · d) / (||q|| · ||d||)
It ranges from −1 (opposite directions) to 1 (identical direction). Using the angle rather than the distance ignores magnitude, which is useful because embedding magnitudes often reflect artefacts like text length rather than meaning. A tiny example in 3 dimensions:
- Query
q = [1.0, 0.2, 0.1] - Doc A
a = [0.9, 0.3, 0.0]→ cosine ≈ 0.98 (nearly parallel: very similar) - Doc B
b = [0.1, 0.9, 0.4]→ cosine ≈ 0.28 (mostly different direction: unrelated)
Doc A wins by a wide margin. At query time, you embed the user's question with the same model, score it against every stored vector, and take the top k. Conceptually simple — the engineering challenge is doing it at scale.
Why a special database?
Brute-force search costs O(n · d): for every query, compare against all n vectors of dimension d. With a million 1536-dimensional vectors, that is 1.5 billion multiply-adds per query — too slow for an interactive product, and it only gets worse as the collection grows. You could also ask: why not a normal database? Because SQL indexes are built for exact matches and range queries on scalars (“price < 50”), not for “find the 10 nearest points in 1536-dimensional space.”
Vector databases exist to answer nearest-neighbour queries approximately but very fast, while handling the boring-but-critical database work: persistence, concurrent writes, metadata filtering (“only search documents from 2024”), and sharding across machines. The “approximate” part is the key algorithmic idea, and it is worth understanding.
Approximate search: IVF and HNSW
Approximate Nearest Neighbour (ANN) search trades a small, controllable amount of accuracy for orders of magnitude of speed. If brute force finds the true top 10, ANN finds 9 or 10 of them in a hundredth of the time. Two index families dominate:
- IVF (Inverted File Index). Cluster all vectors into, say, 1,000 groups with k-means. At query time, find the few clusters nearest the query and search only inside them. If your data splits into 1,000 clusters and you probe 10, you skip ~99% of the comparisons. The parameter
nprobe(how many clusters to check) is the dial between speed and recall. - HNSW (Hierarchical Navigable Small World). Build a layered graph where each vector links to its neighbours, with sparse long-range links on upper layers and dense local links below. A query enters at the top, greedily hops to the closest neighbour at each layer, and descends — like zooming from a country map to a street map. It is currently the best speed-recall trade-off for most workloads and the default in Pinecone, Weaviate, Qdrant, and pgvector's HNSW mode.
The honest trade-off: ANN can miss a true neighbour. In practice, tuned indexes recover 95–99% of the exact top-k while running 10–100× faster, and for RAG that is plenty — the language model is robust to the occasional slightly-suboptimal document.
The RAG pipeline, end to end
RAG has two phases. Indexing happens once (and incrementally as documents change):
- Chunk documents into passages of a few hundred tokens each. Embeddings degrade on very long texts, and retrieval works best at paragraph granularity.
- Embed each chunk with the embedding model and store the vectors in the database alongside the original text and metadata (source, date, page).
Querying happens per question:
- Embed the user's question with the same embedding model.
- Retrieve the top k chunks (typically 3–10) by vector similarity, optionally filtered by metadata.
- Stuff the retrieved text into the prompt ahead of the question, with an instruction like “answer using only the provided context.”
- Generate. The model now answers from your documents, and you can show the sources it used.
Notice what changed versus plain generation: the facts are no longer a property of the model's weights. Update a document, re-embed its chunks, and the model's answers change immediately — no retraining.
Chunking: the detail everyone underestimates
Beginners obsess over which vector database to use; practitioners obsess over chunking. Retrieve too-coarse chunks and you waste context window on irrelevant text; too fine and each chunk lacks the context to be useful or findable. Common strategies:
- Fixed-size with overlap — e.g., 500 tokens with 50 tokens of overlap so ideas spanning a boundary survive. Simple, surprisingly effective.
- Recursive splitting — split on paragraphs, then sentences, then words, keeping chunks under the size limit. Respects document structure.
- Semantic chunking — split where the topic shifts, detected by drops in embedding similarity between adjacent sentences. More expensive, better boundaries.
Attach metadata to every chunk — document title, section heading, date. A surprising amount of retrieval quality comes not from the vector math but from filtering (“only the 2024 pricing docs”) before the similarity search runs.
Where RAG breaks
RAG is not magic, and its failure modes are worth knowing before you build on it:
- Retrieval misses. If the right document is not in the top k, the model cannot use it — and may hallucinate instead of saying “not found.” Hybrid search (vectors plus classic keyword/BM25 scoring) reliably beats pure vector search here.
- Garbage in, garbage out. Contradictory or outdated documents produce contradictory answers. RAG inherits your corpus's quality.
- Lost in the middle. Models attend best to the start and end of long contexts; key facts buried mid-context get underused. Keep k modest and re-rank.
- Evaluation is hard. “Did the answer use the retrieved facts correctly?” needs its own test set and judges — human or LLM-based — not just vibes.
Putting it together
The full picture is a loop: text becomes vectors, vectors become geometry, geometry becomes retrieval, retrieval becomes context, context becomes answers grounded in your data. The vector database is the piece that makes the geometry queryable at production speed — everything else is plumbing around it. If you understand embeddings and approximate nearest-neighbour search, you understand 80% of every RAG system ever built; the rest is chunking strategy and prompt engineering.
Further reading
- Lewis, P., et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020. arXiv:2005.11401
- Johnson, J., Douze, M. & Jégou, H. (2019). Billion-scale similarity search with GPUs. IEEE Trans. Big Data. arXiv:1702.08734 — the FAISS paper; the standard reference on ANN index design.
- Malkov, Y. A. & Yashunin, D. A. (2018). Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs. arXiv:1603.09320
- Gao, Y., et al. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997






