Keyword Search vs Semantic Search
Keyword search vs semantic search is a comparison between two retrieval approaches: matching the literal terms of a query against an index, and matching the meaning of a query through embedding vectors. Keyword search excels at exact identifiers, whereas semantic search excels at paraphrased and descriptive queries.
- Keyword search: Also called lexical or full-text search, it retrieves documents that contain the query terms.
- Inverted index: The data structure behind keyword search, mapping every term to the documents that contain it.
- BM25: A ranking function that scores term matches by frequency, rarity across the collection and document length.
- Semantic search: It converts text into embeddings and ranks documents by vector similarity.
- Vocabulary mismatch: The situation where the searcher and the author describe the same problem in different words.
For example, the query "ECONNREFUSED 5432" is answered precisely by keyword search, while the query "app cannot reach the db" is answered better by semantic search over the same engineering runbooks.
Quick Answer
Keyword search is the better choice when queries contain exact tokens such as error codes, stack trace fragments, configuration keys or function names. Semantic search is the better choice when users describe symptoms in natural language that differs from the documentation. Most engineering knowledge bases need both, which is the purpose of hybrid search.
Keyword Search vs Semantic Search: Comparison Table
| Aspect | Keyword Search | Semantic Search |
|---|---|---|
| Matching unit | Literal terms and tokens | Meaning encoded in vectors |
| Index structure | Inverted index of terms | Vector index, such as HNSW |
| Typical scoring | BM25 or TF-IDF | Cosine similarity or dot product |
| Exact identifiers | Precise, including rare codes | Frequently imprecise |
| Paraphrased queries | Misses documents with other wording | Retrieves them reliably |
| Explainability | Matched terms can be highlighted | Similarity scores are opaque |
| Indexing cost | Tokenisation only | Embedding computation for every chunk |
| Model dependence | None, apart from analysers | Depends on the embedding model |
| Typical failure | Vocabulary mismatch | Topical drift and lost precision |
Scoring with TF-IDF and BM25
Keyword search converts term matches into a numeric relevance score. The BM25 formula descends from TF-IDF, an older weighting scheme that multiplies how frequently a term occurs in a document by how rare that term is across the whole collection:
BM25 keeps both ingredients, so the TF-IDF vs BM25 difference comes from two corrections: repetition has a ceiling, and longer documents need additional repetitions to earn the same credit.
- : the query, treated as a collection of individual terms .
- : one document, and is its length measured in terms.
- : the term frequency, meaning how many times the term occurs in the document.
- : the total number of documents in the indexed collection.
- : the document frequency, meaning how many documents contain the term at least once.
- : the average document length across the collection.
- : the saturation parameter, usually between 1.2 and 2.0, which limits the benefit of repeating a term.
- : the length normalisation parameter, usually 0.75, where 0 ignores length entirely and 1 applies full normalisation.
In words, every query term contributes its rarity weight multiplied by a frequency factor. The frequency factor increases with every repetition but can never exceed , however often the term appears. The additional 1 inside the logarithm, used by Lucene and many other search engines, keeps the rarity weight positive even for extremely common terms.
The worked example scores three short snippets for the query "504 timeout", using the standard settings and :
- Collection statistics: The three snippets contain 5, 6 and 4 terms, so the average document length equals 5.
- Rarity weights: The status code 504 appears in a single snippet, so it receives a considerably larger IDF than timeout, which appears in two snippets.
- runbook-504: Its length equals the average, so the length correction equals 1, and the two occurrences of timeout are partially discounted by saturation.
- api-timeouts: Its six terms exceed the average length, so the length correction rises to 1.15, and its single timeout earns slightly less than the full IDF.
- runbook-disk: None of the query terms occurs in this snippet, so its relevance score remains exactly 0.
The program below repeats the calculation and prints the frequency factor for an average-length document.
import math
# Three tiny snippets from an engineering knowledge base.
DOCS = {
"runbook-504": "504 gateway timeout upstream timeout",
"api-timeouts": "request timeout defaults to 30 seconds",
"runbook-disk": "disk full rotate logs",
}
QUERY = "504 timeout"
K1, B = 1.2, 0.75
docs = {name: text.split() for name, text in DOCS.items()}
N = len(docs)
avgdl = sum(len(t) for t in docs.values()) / N
print(f"N = {N}, avgdl = {avgdl:.1f}")
def idf(term):
n = sum(term in t for t in docs.values()) # documents containing the term
return math.log(1 + (N - n + 0.5) / (n + 0.5))
for term in QUERY.split():
n = sum(term in t for t in docs.values())
print(f"{term:<8} n = {n} TF-IDF idf = {math.log(N / n):.3f} BM25 IDF = {idf(term):.3f}")
for name, terms in docs.items():
total = 0.0
parts = []
for term in QUERY.split():
f = terms.count(term)
norm = 1 - B + B * len(terms) / avgdl
part = idf(term) * f * (K1 + 1) / (f + K1 * norm)
total += part
parts.append(f"{term} {part:.3f}")
print(f"{name:<13} |D| = {len(terms)} " + ", ".join(parts) + f" score = {total:.3f}")
# Term-frequency saturation for an average-length document (|D| = avgdl).
print("f: " + " ".join(f"{f:>5}" for f in (1, 2, 3, 5, 10)))
print("tf: " + " ".join(f"{f * (K1 + 1) / (f + K1):5.2f}" for f in (1, 2, 3, 5, 10)))N = 3, avgdl = 5.0
504 n = 1 TF-IDF idf = 1.099 BM25 IDF = 0.981
timeout n = 2 TF-IDF idf = 0.405 BM25 IDF = 0.470
runbook-504 |D| = 5 504 0.981, timeout 0.646 score = 1.627
api-timeouts |D| = 6 504 0.000, timeout 0.434 score = 0.434
runbook-disk |D| = 4 504 0.000, timeout 0.000 score = 0.000
f: 1 2 3 5 10
tf: 1.00 1.38 1.57 1.77 1.96- Rarity dominates: The single rare status code contributes more to the total than two occurrences of the common word timeout.
- Saturation: A second occurrence raises the frequency factor from 1.00 to only 1.38, whereas plain TF-IDF would double that term's contribution.
- Length penalty: The api-timeouts snippet is one term longer than average, so its single match earns 0.434 instead of the full 0.470.
In practice, and are the principal tuning controls of a keyword index, and adjusting them prevents long API reference pages from outranking short, precise runbooks.
When to Use Keyword Search
- Error messages: Queries such as "ECONNREFUSED" or "OOMKilled" must match the exact token in logs and runbooks.
- Code identifiers: Function names, class names and configuration keys have no meaningful paraphrase.
- Version-specific content: Queries that include a version number or ticket identifier depend on exact matches.
- Transparent ranking: Teams that need to explain and debug rankings benefit from visible term matches.
- Limited infrastructure: Keyword search requires no embedding model and no vector index.
When to Use Semantic Search
- Symptom descriptions: Engineers describe what they observe, not the exact wording of the runbook title.
- Conceptual questions: Questions about how a system behaves rarely repeat the vocabulary of the documentation.
- Cross-team terminology: Different teams use different names for the same component or failure.
- Retrieval for language models: Retrieval-augmented generation depends on finding relevant passages for natural-language questions.
How Each Method Scores a Document
- Term frequency: BM25 rewards documents that repeat a query term, with diminishing returns controlled by the parameter k1.
- Inverse document frequency: Rare terms such as a specific port number contribute far more to the score than common words.
- Length normalisation: The parameter b penalises long documents, so a short runbook that mentions a term competes fairly with a long one.
- Vector similarity: Semantic search computes the cosine of the angle between the query vector and each document vector.
- Score scales: BM25 scores are unbounded and collection-dependent, whereas cosine similarity is bounded between -1 and 1, so the two scores cannot be added directly.
Example: Runbook Search with BM25 and Embeddings
The program below compares BM25 vs embeddings on four runbook snippets: a complete BM25 implementation against a ranking from hand-made meaning vectors.
import math
DOCS = {
"pg-refused": "error ECONNREFUSED 127.0.0.1:5432 means postgres is down or the port is blocked",
"redis-refused": "error ECONNREFUSED 127.0.0.1:6379 means the redis cache is not running",
"pool-exhausted": "the service cannot get a database connection because the pool is exhausted",
"jwt-expired": "401 token expired: refresh the JWT before calling the API",
}
# Hand-made meaning vectors: [database, network, cache, auth]
DOC_VECS = {
"pg-refused": [0.8, 0.6, 0.0, 0.0], "redis-refused": [0.1, 0.6, 0.8, 0.0],
"pool-exhausted": [0.9, 0.3, 0.0, 0.0], "jwt-expired": [0.0, 0.1, 0.0, 1.0],
}
QUERIES = {
"ECONNREFUSED 5432": [0.4, 0.8, 0.4, 0.0],
"app cannot reach the db": [0.9, 0.4, 0.0, 0.0],
}
STOP = {"the", "is", "or", "a"}
def tokens(text):
return [w for w in text.lower().replace(":", " ").split() if w not in STOP]
def bm25_rank(query, k1=1.5, b=0.75):
docs = {d: tokens(t) for d, t in DOCS.items()}
avgdl = sum(len(t) for t in docs.values()) / len(docs)
scores = {}
for d, terms in docs.items():
score = 0.0
for term in tokens(query):
df = sum(term in t for t in docs.values())
if df == 0:
continue
idf = math.log(1 + (len(docs) - df + 0.5) / (df + 0.5))
tf = terms.count(term)
score += idf * tf * (k1 + 1) / (tf + k1 * (1 - b + b * len(terms) / avgdl))
scores[d] = score
return sorted(scores.items(), key=lambda x: -x[1])
def cosine(a, b):
dot = sum(x * y for x, y in zip(a, b))
return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))
for query, qvec in QUERIES.items():
print(f'Query: "{query}"')
keyword = [f"{d} {s:.2f}" for d, s in bm25_rank(query)[:2] if s > 0]
semantic = sorted(DOC_VECS, key=lambda d: -cosine(qvec, DOC_VECS[d]))[:2]
print(" keyword (BM25): ", ", ".join(keyword))
print(" semantic: ", ", ".join(f"{d} {cosine(qvec, DOC_VECS[d]):.2f}" for d in semantic))Query: "ECONNREFUSED 5432"
keyword (BM25): pg-refused 1.85, redis-refused 0.68
semantic: redis-refused 0.85, pg-refused 0.82
Query: "app cannot reach the db"
keyword (BM25): pool-exhausted 1.24
semantic: pool-exhausted 1.00, pg-refused 0.97- Exact token query: BM25 ranks pg-refused first because the rare token 5432 appears only in that document, while the vectors confuse the two refused-connection runbooks.
- Descriptive query: BM25 finds pool-exhausted only through the incidental word "cannot" and misses pg-refused entirely, whereas semantic search retrieves both relevant runbooks.
- Complementary errors: Each method fails on the query that the other method handles well, which motivates combining them.
Combining Both Methods
In production, lexical search vs semantic search is rarely an either-or choice. Hybrid search runs both retrievers and merges their rankings, and a reranking model, described in reranking in RAG, can then reorder the merged candidates. The nearest-neighbour step inside semantic search is covered in vector search.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In the Python example, why does semantic search rank redis-refused above pg-refused for "ECONNREFUSED 5432"?
Frequently Asked Questions
What is the difference between keyword and semantic search?
Keyword search matches the exact terms of a query against an inverted index and scores them with a formula such as BM25. Semantic search converts the query and documents into embedding vectors and ranks documents by vector similarity, so it matches meaning rather than wording.
Is BM25 a keyword search algorithm?
Yes. BM25 is a ranking function for lexical search that scores documents by term frequency, the rarity of each term across the collection and document length. It is the default scoring method in many full-text search engines.
Is semantic search always more accurate than keyword search?
No. Semantic search performs better on paraphrased and descriptive queries, but keyword search is usually more accurate for exact identifiers such as error codes, function names and ticket numbers. The better choice depends on the queries users actually type.
Can keyword search and semantic search be combined?
Yes. Hybrid search runs both methods on the same query and merges their rankings, commonly with reciprocal rank fusion. This covers both exact terms and paraphrased intent.
Related Articles
- What is Semantic SearchSemantic search explained: how embeddings match meaning instead of exact words, how the pipeline works, where it fails, and a Python runbook search demo.
- Hybrid Search (BM25 + Vector)Hybrid search explained: how BM25 keyword results and vector results are merged with reciprocal rank fusion, with a runnable Python runbook search example.
- Vector Search ExplainedVector search explained: exact k-nearest-neighbour search, ANN indexes such as HNSW and IVF, similarity metrics, and a brute-force Python code search demo.
- Embeddings in LLMLearn what embeddings in LLMs are, how text becomes vectors, how cosine similarity compares meaning and where embeddings are used, with Python code.
- Reranking in RAGReranking in RAG explained: how a second, stricter scorer reorders retrieved passages, cross-encoders vs bi-encoders, a Python example and the trade-offs.
- How AI Search Engines Work (Perplexity, AI Overviews)How AI search engines work: query rewriting, retrieval, passage extraction and cited answer generation, with a Python pipeline over engineering docs.