What is Semantic Search
Semantic search is a retrieval method that finds documents by the meaning of a query instead of its exact words. It converts the query and every document into embeddings, numerical vectors that encode meaning, and returns the documents whose vectors lie closest to the query vector.
- Embedding model: A neural network that maps a passage of text to a fixed-length numerical representation.
- Vector: An ordered list of numbers in which semantically related passages produce nearby positions.
- Similarity: A numerical score, usually cosine similarity, that measures the proximity of two vectors.
- Index: A collection of document vectors, frequently maintained in a vector database.
- Intent: The underlying information need behind a query, which users can express in many different formulations.
For example, an engineer who searches the team knowledge base for "app cannot reach the db" receives the runbook titled "Postgres connection refused on port 5432", although the two texts share no words.
Key Characteristics of Semantic Search
- Meaning over wording: As a meaning-based search method, it retrieves the same relevant document for synonyms, paraphrases and alternative terminology.
- Dense representation: Every document becomes a vector with hundreds of dimensions instead of a sparse list of terms.
- Ranked results: Documents are returned in descending order of similarity, each accompanied by a relevance score.
- Vocabulary independence: Informal descriptions from engineers can match the formal terminology used in API documentation.
- Model dependence: Retrieval quality depends heavily on the embedding model, its training data and its dimensionality.
- Limited precision: Exact tokens such as error codes, port numbers and version identifiers carry relatively little weight.
How Semantic Search Works
- Chunking: Documents such as runbooks, API references and postmortems are divided into passages of a few hundred words.
- Embedding: The embedding model converts each passage into a vector, which is stored together with its text, source and metadata.
- Query encoding: The query is embedded with the identical model, so the query vector and the document vectors occupy one space.
- Similarity search: The system identifies the stored vectors nearest to the query vector, a procedure explained in vector search.
- Ranking: The candidate documents are sorted by similarity and are optionally reordered by a separate reranking model.
- Presentation: The results appear as links with snippets, or they are supplied to a language model as retrieved context.
Example: Semantic Search over Runbooks in Python
The semantic search example below builds toy embeddings from a hand-made concept lexicon and ranks three runbook titles by cosine similarity.
import math
# Hand-made concept vectors: [database, network, auth, memory]
LEXICON = {
"postgres": [1, 0, 0, 0], "database": [1, 0, 0, 0], "db": [1, 0, 0, 0],
"connection": [0.5, 0.5, 0, 0], "refused": [0, 1, 0, 0],
"reach": [0, 1, 0, 0], "port": [0, 0.8, 0, 0],
"token": [0, 0, 1, 0], "expired": [0, 0, 1, 0], "login": [0, 0, 1, 0],
"heap": [0, 0, 0, 1], "oom": [0, 0, 0, 1], "killed": [0, 0.2, 0, 0.8],
}
def embed(text):
vec = [0.0] * 4
for word in text.lower().split():
for i, x in enumerate(LEXICON.get(word, [0, 0, 0, 0])):
vec[i] += x
return vec
def cosine(a, b):
dot = sum(x * y for x, y in zip(a, b))
norm = math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b))
return dot / norm if norm else 0.0
DOCS = {
"runbook/postgres-connection-refused": "Postgres connection refused on port 5432",
"runbook/jwt-token-expired": "JWT token expired during login",
"runbook/pod-oom-killed": "Pod OOM killed after heap growth",
}
query = "app cannot reach the db"
print("Query:", query)
print("Shared words with each title:")
for doc_id, text in DOCS.items():
shared = set(query.split()) & set(text.lower().split())
print(f" {doc_id:<37} {sorted(shared)}")
print("Semantic ranking (cosine similarity):")
q = embed(query)
ranked = sorted(DOCS, key=lambda d: cosine(q, embed(DOCS[d])), reverse=True)
for doc_id in ranked:
print(f" {doc_id:<37} {cosine(q, embed(DOCS[doc_id])):.3f}")Query: app cannot reach the db
Shared words with each title:
runbook/postgres-connection-refused []
runbook/jwt-token-expired []
runbook/pod-oom-killed []
Semantic ranking (cosine similarity):
runbook/postgres-connection-refused 0.979
runbook/pod-oom-killed 0.050
runbook/jwt-token-expired 0.000- No shared vocabulary: A literal word match returns nothing, because the query and every title have zero terms in common.
- Shared meaning: The words "db" and "reach" map to the database and network concepts, which the postgres runbook also expresses.
- Simplification: A production embedding model learns these relationships from enormous text corpora instead of a handwritten lexicon.
Applications of Semantic Search
- Knowledge base search: Retrieving runbooks and postmortems that describe an incident in different terminology.
- Code search: Locating functions by their behaviour, such as "retry with exponential backoff".
- Ticket deduplication: Identifying bug reports that describe the identical failure in different language.
- Documentation assistants: Supplying relevant passages to retrieval-augmented generation (RAG).
- Question answering: Matching a natural-language question to the most similar existing FAQ entry.
- AI search engines: Operating as the retrieval stage described in how AI search engines work.
Advantages
- Recall on paraphrases: Users do not need to anticipate the exact vocabulary chosen by the original author.
- Multilingual retrieval: Multilingual embedding models can match a query in one language to a document in another.
- Reduced maintenance: It requires no manually curated synonym dictionaries or stemming configuration.
- Natural queries: Complete questions and descriptive error reports work as effectively as short keyword queries.
Limitations
- Exact identifiers: Error codes, port numbers and function names are often ranked poorly, as shown in keyword search vs semantic search.
- Opaque scores: A similarity score provides no explanation of why a particular document matched.
- Computation cost: Every document must be embedded initially and embedded again whenever the model changes.
- Topical drift: A document on a related topic can outrank the specific document that answers the question.
- Domain gaps: General-purpose embedding models may not understand internal service names and organisational terminology.
Most production systems pair semantic search with keyword matching in hybrid search to cover both meaning and exact terms.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In the Python example, why does the postgres runbook rank first for the query "app cannot reach the db"?
Frequently Asked Questions
What is semantic search in simple terms?
Semantic search finds documents by what a query means rather than by the exact words it contains. It converts the query and the documents into embedding vectors and returns the documents whose vectors are closest to the query vector.
Is semantic search the same as vector search?
Not exactly. Semantic search is the goal of matching by meaning, while vector search is the retrieval technique that finds the nearest vectors. Most semantic search systems use vector search as their core step.
Does semantic search replace keyword search?
Usually not. Keyword search remains more reliable for exact identifiers such as error codes, function names and version numbers. Many production systems combine both methods in a hybrid search.
What is needed to build semantic search with embeddings?
It needs an embedding model, a store for the document vectors and a similarity search step. A reranker and metadata filters are common additions that improve the final ranking.
Related Articles
- Keyword Search vs Semantic SearchKeyword search vs semantic search compared: BM25 term matching versus embedding similarity, a comparison table, when to use each, and a Python BM25 demo.
- Vector Search ExplainedVector search explained: exact k-nearest-neighbour search, ANN indexes such as HNSW and IVF, similarity metrics, and a brute-force Python code search demo.
- Hybrid Search (BM25 + Vector)Hybrid search explained: how BM25 keyword results and vector results are merged with reciprocal rank fusion, with a runnable Python runbook search example.
- Embeddings in LLMLearn what embeddings in LLMs are, how text becomes vectors, how cosine similarity compares meaning and where embeddings are used, with Python code.
- Vector Database in RAGVector database in RAG explained: how embeddings are stored, indexed and searched with cosine similarity, ANN indexes, metadata filters and a Python demo.
- What is RAG (Retrieval-Augmented Generation)Retrieval-augmented generation (RAG) explained: how it retrieves document passages to ground LLM answers, its key steps, a Python example and its limits.