Vector Database in RAG
A vector database is a data store that saves embeddings, numerical vectors that represent meaning, together with their source text and metadata. In a RAG system, it performs similarity search, returning the stored chunks whose vectors lie closest to the vector of a user question.
- Vector: A fixed-length list of numbers produced by an embedding model for one chunk of text.
- Metric: The function that measures closeness between vectors, commonly cosine similarity, dot product or Euclidean distance.
- Top-k: A query that returns the k most similar vectors, ranked by score.
- Filter: A condition on stored fields, such as version or product, applied during the search.
- ANN index: An approximate nearest neighbour structure that searches millions of vectors without comparing every one.
For example, a documentation assistant stores one vector per Orders API section and, for the question "Why do I get 429 Too Many Requests?", returns the Rate limits and Error codes chunks.
Key Characteristics of a Vector Database
- Similarity: Records are retrieved by distance in vector space instead of exact key or keyword matches.
- Dimensionality: Every vector in a collection has the same dimension, determined by the embedding model.
- Approximation: Index types such as HNSW and IVF trade a small loss of recall for much lower query latency.
- Metadata: Each vector is stored with its chunk text, source URL, section and version for filtering and citation.
- Mutability: Individual vectors can be inserted, updated or deleted when the underlying documents change.
- Scalability: Collections can be sharded and replicated across nodes to serve high query volumes.
How a Vector Database Works
- Ingestion: The application writes each chunk's vector, identifier, text and metadata into a collection.
- Normalisation: For cosine similarity, vectors are often scaled to unit length so that a dot product yields the cosine.
- Indexing: The database builds an approximate search structure, such as an HNSW index (a Hierarchical Navigable Small World graph), that links each vector to its near neighbours.
- Querying: The question is embedded with the same model, and the database traverses the index towards the closest vectors.
- Filtering: Metadata conditions remove ineligible records, such as documentation for a retired API version.
- Ranking: The remaining candidates are sorted by similarity, and the top-k records are returned with their scores.
Example: An In-Memory Vector Store in Python
The code below implements a minimal vector store with exact cosine similarity search, top-k results and a metadata filter, using hand-made four-dimensional vectors.
import math
class VectorStore:
"""A tiny in-memory vector store: exact (brute-force) cosine search."""
def __init__(self, dim):
self.dim, self.rows = dim, []
def add(self, chunk_id, vector, metadata):
assert len(vector) == self.dim, "vector has the wrong dimension"
norm = math.sqrt(sum(x * x for x in vector))
self.rows.append((chunk_id, [x / norm for x in vector], metadata))
def search(self, query, k=3, where=None):
norm = math.sqrt(sum(x * x for x in query))
q = [x / norm for x in query]
hits = []
for chunk_id, vec, meta in self.rows:
if where and any(meta.get(key) != val for key, val in where.items()):
continue # metadata filter runs before scoring
hits.append((sum(a * b for a, b in zip(q, vec)), chunk_id))
return sorted(hits, reverse=True)[:k]
# Hand-made 4-dimensional "embeddings": [auth, limits, paging, events]
store = VectorStore(dim=4)
store.add("authentication", [0.9, 0.1, 0.0, 0.1], {"version": "v2"})
store.add("rate-limits", [0.1, 0.9, 0.2, 0.0], {"version": "v2"})
store.add("pagination", [0.0, 0.3, 0.9, 0.1], {"version": "v2"})
store.add("webhooks", [0.2, 0.0, 0.1, 0.9], {"version": "v2"})
store.add("error-codes", [0.4, 0.6, 0.1, 0.2], {"version": "v2"})
store.add("rate-limits-v1", [0.1, 0.9, 0.1, 0.0], {"version": "v1"})
query = [0.2, 0.9, 0.1, 0.1] # "Why do I get 429 Too Many Requests?"
print("Top 3, all versions:")
for score, chunk_id in store.search(query, k=3):
print(f" {chunk_id:<15} {score:.3f}")
print("Top 3, version v2 only:")
for score, chunk_id in store.search(query, k=3, where={"version": "v2"}):
print(f" {chunk_id:<15} {score:.3f}")Output:
Top 3, all versions:
rate-limits-v1 0.989
rate-limits 0.983
error-codes 0.923
Top 3, version v2 only:
rate-limits 0.983
error-codes 0.923
pagination 0.416- Staleness: Without a filter, the outdated v1 rate-limit chunk ranks first, because similarity alone ignores document versions.
- Filtering: The version filter removes the v1 chunk before scoring, so the current documentation ranks first.
- Complexity: This store compares the query with every vector, which is accurate but scales linearly; production databases use ANN indexes instead.
Applications of Vector Databases
- RAG retrieval: Serving as the vector database for RAG pipelines and supplying relevant chunks to a language model, as described in how RAG works.
- Semantic search: Powering semantic search over documentation, tickets and runbooks.
- Code search: Finding functions with similar behaviour across large repositories.
- Deduplication: Detecting near-duplicate bug reports or support tickets.
- Recommendation: Suggesting related articles or API examples by vector proximity.
- Agent memory: Storing past interactions that an agent can recall by similarity.
Advantages
- Relevance: Paraphrased questions still match relevant chunks, unlike exact keyword lookup.
- Latency: ANN indexes return results from millions of vectors within milliseconds.
- Expressiveness: Metadata filters and similarity ranking run in a single query.
- Maintainability: Changed chunks can be replaced without rebuilding the language model.
Limitations
- Recall: ANN search can miss a true nearest neighbour, especially with aggressive index settings.
- Identifiers: Exact identifiers such as error code 429 or next_cursor can rank poorly, which hybrid search addresses.
- Migration: Changing the embedding model requires regenerating and re-indexing every vector.
- Memory footprint: HNSW graphs are usually held in memory, which raises infrastructure cost at large scale.
- Ranking: The closest vectors are not always the most useful passages, so many systems add reranking in RAG.
Common options include PostgreSQL with the pgvector extension, Pinecone, Weaviate, Qdrant, Milvus and Chroma. The underlying search algorithms are covered in vector search explained.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In the Python example, why does rate-limits-v1 rank first when no filter is applied?
Frequently Asked Questions
What is a vector database used for in RAG?
A vector database stores the embedding of every document chunk together with its text and metadata. At query time it returns the chunks whose vectors are most similar to the question vector, and those chunks become the context for the language model.
Is a vector database different from a vector index?
A vector index is the data structure that makes similarity search fast, such as HNSW. A vector database wraps one or more indexes with storage, metadata filtering, updates, access control and replication.
Can PostgreSQL be used as a vector database?
Yes, through an extension that adds a vector column type and similarity operators. This suits teams that already run PostgreSQL and want chunk vectors next to their relational data.
Which similarity metric should a vector database use?
The metric should match the one the embedding model was trained for, which is usually cosine similarity or dot product. For normalised vectors, cosine similarity and dot product produce the same ranking.
Related Articles
- How RAG WorksHow RAG works in detail: the offline indexing pipeline, the online query pipeline, a traced Python example over API docs, and where each stage can fail.
- Chunking Strategies in RAGChunking strategies in RAG explained: fixed-size, heading-based and semantic chunking, how to choose chunk size and overlap, with a Python comparison.
- Embeddings in LLMLearn what embeddings in LLMs are, how text becomes vectors, how cosine similarity compares meaning and where embeddings are used, with Python code.
- Vector Search ExplainedVector search explained: exact k-nearest-neighbour search, ANN indexes such as HNSW and IVF, similarity metrics, and a brute-force Python code search demo.
- Hybrid Search (BM25 + Vector)Hybrid search explained: how BM25 keyword results and vector results are merged with reciprocal rank fusion, with a runnable Python runbook search example.
- Reranking in RAGReranking in RAG explained: how a second, stricter scorer reorders retrieved passages, cross-encoders vs bi-encoders, a Python example and the trade-offs.