…
Skip to content
Topics
On this page

Keyword Search vs Semantic Search

Keyword search vs semantic search is a comparison between two retrieval approaches: matching the literal terms of a query against an index, and matching the meaning of a query through embedding vectors. Keyword search excels at exact identifiers, whereas semantic search excels at paraphrased and descriptive queries.

  • Keyword search: Also called lexical or full-text search, it retrieves documents that contain the query terms.
  • Inverted index: The data structure behind keyword search, mapping every term to the documents that contain it.
  • BM25: A ranking function that scores term matches by frequency, rarity across the collection and document length.
  • Semantic search: It converts text into embeddings and ranks documents by vector similarity.
  • Vocabulary mismatch: The situation where the searcher and the author describe the same problem in different words.
Keyword search vs semantic search on the same runbooksTwo lanes over the same engineering runbooks. In the top lane, the query ECONNREFUSED 5432 goes to keyword search with BM25, which matches exact tokens and returns pg-refused. In the bottom lane, the query "app cannot reach the db" goes to semantic search, which compares vectors and returns pool-exhausted and pg-refused. Results are illustrative.ECONNREFUSED 5432keyword (BM25)exact tokenspg-refusedapp cannot reach the dbsemanticvector similaritypool-exhaustedpg-refused
Keyword search vs semantic search on the same runbooks

For example, the query "ECONNREFUSED 5432" is answered precisely by keyword search, while the query "app cannot reach the db" is answered better by semantic search over the same engineering runbooks.

Quick Answer

Keyword search is the better choice when queries contain exact tokens such as error codes, stack trace fragments, configuration keys or function names. Semantic search is the better choice when users describe symptoms in natural language that differs from the documentation. Most engineering knowledge bases need both, which is the purpose of hybrid search.

Keyword Search vs Semantic Search: Comparison Table

AspectKeyword SearchSemantic Search
Matching unitLiteral terms and tokensMeaning encoded in vectors
Index structureInverted index of termsVector index, such as HNSW
Typical scoringBM25 or TF-IDFCosine similarity or dot product
Exact identifiersPrecise, including rare codesFrequently imprecise
Paraphrased queriesMisses documents with other wordingRetrieves them reliably
ExplainabilityMatched terms can be highlightedSimilarity scores are opaque
Indexing costTokenisation onlyEmbedding computation for every chunk
Model dependenceNone, apart from analysersDepends on the embedding model
Typical failureVocabulary mismatchTopical drift and lost precision

Scoring with TF-IDF and BM25

Keyword search converts term matches into a numeric relevance score. The BM25 formula descends from TF-IDF, an older weighting scheme that multiplies how frequently a term occurs in a document by how rare that term is across the whole collection:

BM25 keeps both ingredients, so the TF-IDF vs BM25 difference comes from two corrections: repetition has a ceiling, and longer documents need additional repetitions to earn the same credit.

  • : the query, treated as a collection of individual terms .
  • : one document, and is its length measured in terms.
  • : the term frequency, meaning how many times the term occurs in the document.
  • : the total number of documents in the indexed collection.
  • : the document frequency, meaning how many documents contain the term at least once.
  • : the average document length across the collection.
  • : the saturation parameter, usually between 1.2 and 2.0, which limits the benefit of repeating a term.
  • : the length normalisation parameter, usually 0.75, where 0 ignores length entirely and 1 applies full normalisation.

In words, every query term contributes its rarity weight multiplied by a frequency factor. The frequency factor increases with every repetition but can never exceed , however often the term appears. The additional 1 inside the logarithm, used by Lucene and many other search engines, keeps the rarity weight positive even for extremely common terms.

The worked example scores three short snippets for the query "504 timeout", using the standard settings and :

  • Collection statistics: The three snippets contain 5, 6 and 4 terms, so the average document length equals 5.
  • Rarity weights: The status code 504 appears in a single snippet, so it receives a considerably larger IDF than timeout, which appears in two snippets.
  • runbook-504: Its length equals the average, so the length correction equals 1, and the two occurrences of timeout are partially discounted by saturation.
  • api-timeouts: Its six terms exceed the average length, so the length correction rises to 1.15, and its single timeout earns slightly less than the full IDF.
  • runbook-disk: None of the query terms occurs in this snippet, so its relevance score remains exactly 0.

The program below repeats the calculation and prints the frequency factor for an average-length document.

Python
import math

# Three tiny snippets from an engineering knowledge base.
DOCS = {
    "runbook-504": "504 gateway timeout upstream timeout",
    "api-timeouts": "request timeout defaults to 30 seconds",
    "runbook-disk": "disk full rotate logs",
}
QUERY = "504 timeout"
K1, B = 1.2, 0.75

docs = {name: text.split() for name, text in DOCS.items()}
N = len(docs)
avgdl = sum(len(t) for t in docs.values()) / N
print(f"N = {N}, avgdl = {avgdl:.1f}")

def idf(term):
    n = sum(term in t for t in docs.values())  # documents containing the term
    return math.log(1 + (N - n + 0.5) / (n + 0.5))

for term in QUERY.split():
    n = sum(term in t for t in docs.values())
    print(f"{term:<8} n = {n}  TF-IDF idf = {math.log(N / n):.3f}  BM25 IDF = {idf(term):.3f}")

for name, terms in docs.items():
    total = 0.0
    parts = []
    for term in QUERY.split():
        f = terms.count(term)
        norm = 1 - B + B * len(terms) / avgdl
        part = idf(term) * f * (K1 + 1) / (f + K1 * norm)
        total += part
        parts.append(f"{term} {part:.3f}")
    print(f"{name:<13} |D| = {len(terms)}  " + ", ".join(parts) + f"  score = {total:.3f}")

# Term-frequency saturation for an average-length document (|D| = avgdl).
print("f:  " + "  ".join(f"{f:>5}" for f in (1, 2, 3, 5, 10)))
print("tf: " + "  ".join(f"{f * (K1 + 1) / (f + K1):5.2f}" for f in (1, 2, 3, 5, 10)))
Output
N = 3, avgdl = 5.0
504      n = 1  TF-IDF idf = 1.099  BM25 IDF = 0.981
timeout  n = 2  TF-IDF idf = 0.405  BM25 IDF = 0.470
runbook-504   |D| = 5  504 0.981, timeout 0.646  score = 1.627
api-timeouts  |D| = 6  504 0.000, timeout 0.434  score = 0.434
runbook-disk  |D| = 4  504 0.000, timeout 0.000  score = 0.000
f:      1      2      3      5     10
tf:  1.00   1.38   1.57   1.77   1.96
BM25 term-frequency saturation for k1 = 1.2: extra repetitions of a term add less and lessA line chart of the BM25 frequency factor against the number of times a term occurs in a document, from 0 to 10, with k1 = 1.2 and b = 0.75. For a document of average length the factor is 1.00 at one occurrence, 1.38 at two, 1.57 at three, 1.77 at five and 1.96 at ten, approaching the limit k1 + 1 = 2.2. A document twice the average length sits lower. A raw count, as in TF-IDF, rises in a straight line without a limit. The values match the worked example, and the documents are illustrative.051012term frequency f(t, D)frequency factorlimit k1 + 1 = 2.2raw count (TF-IDF)average lengthtwice average length
BM25 term-frequency saturation for k1 = 1.2: extra repetitions of a term add less and less
  • Rarity dominates: The single rare status code contributes more to the total than two occurrences of the common word timeout.
  • Saturation: A second occurrence raises the frequency factor from 1.00 to only 1.38, whereas plain TF-IDF would double that term's contribution.
  • Length penalty: The api-timeouts snippet is one term longer than average, so its single match earns 0.434 instead of the full 0.470.

In practice, and are the principal tuning controls of a keyword index, and adjusting them prevents long API reference pages from outranking short, precise runbooks.

  • Error messages: Queries such as "ECONNREFUSED" or "OOMKilled" must match the exact token in logs and runbooks.
  • Code identifiers: Function names, class names and configuration keys have no meaningful paraphrase.
  • Version-specific content: Queries that include a version number or ticket identifier depend on exact matches.
  • Transparent ranking: Teams that need to explain and debug rankings benefit from visible term matches.
  • Limited infrastructure: Keyword search requires no embedding model and no vector index.
  • Symptom descriptions: Engineers describe what they observe, not the exact wording of the runbook title.
  • Conceptual questions: Questions about how a system behaves rarely repeat the vocabulary of the documentation.
  • Cross-team terminology: Different teams use different names for the same component or failure.
  • Retrieval for language models: Retrieval-augmented generation depends on finding relevant passages for natural-language questions.

How Each Method Scores a Document

  • Term frequency: BM25 rewards documents that repeat a query term, with diminishing returns controlled by the parameter k1.
  • Inverse document frequency: Rare terms such as a specific port number contribute far more to the score than common words.
  • Length normalisation: The parameter b penalises long documents, so a short runbook that mentions a term competes fairly with a long one.
  • Vector similarity: Semantic search computes the cosine of the angle between the query vector and each document vector.
  • Score scales: BM25 scores are unbounded and collection-dependent, whereas cosine similarity is bounded between -1 and 1, so the two scores cannot be added directly.

Example: Runbook Search with BM25 and Embeddings

The program below compares BM25 vs embeddings on four runbook snippets: a complete BM25 implementation against a ranking from hand-made meaning vectors.

Python
import math

DOCS = {
    "pg-refused": "error ECONNREFUSED 127.0.0.1:5432 means postgres is down or the port is blocked",
    "redis-refused": "error ECONNREFUSED 127.0.0.1:6379 means the redis cache is not running",
    "pool-exhausted": "the service cannot get a database connection because the pool is exhausted",
    "jwt-expired": "401 token expired: refresh the JWT before calling the API",
}
# Hand-made meaning vectors: [database, network, cache, auth]
DOC_VECS = {
    "pg-refused": [0.8, 0.6, 0.0, 0.0], "redis-refused": [0.1, 0.6, 0.8, 0.0],
    "pool-exhausted": [0.9, 0.3, 0.0, 0.0], "jwt-expired": [0.0, 0.1, 0.0, 1.0],
}
QUERIES = {
    "ECONNREFUSED 5432": [0.4, 0.8, 0.4, 0.0],
    "app cannot reach the db": [0.9, 0.4, 0.0, 0.0],
}

STOP = {"the", "is", "or", "a"}

def tokens(text):
    return [w for w in text.lower().replace(":", " ").split() if w not in STOP]

def bm25_rank(query, k1=1.5, b=0.75):
    docs = {d: tokens(t) for d, t in DOCS.items()}
    avgdl = sum(len(t) for t in docs.values()) / len(docs)
    scores = {}
    for d, terms in docs.items():
        score = 0.0
        for term in tokens(query):
            df = sum(term in t for t in docs.values())
            if df == 0:
                continue
            idf = math.log(1 + (len(docs) - df + 0.5) / (df + 0.5))
            tf = terms.count(term)
            score += idf * tf * (k1 + 1) / (tf + k1 * (1 - b + b * len(terms) / avgdl))
        scores[d] = score
    return sorted(scores.items(), key=lambda x: -x[1])

def cosine(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))

for query, qvec in QUERIES.items():
    print(f'Query: "{query}"')
    keyword = [f"{d} {s:.2f}" for d, s in bm25_rank(query)[:2] if s > 0]
    semantic = sorted(DOC_VECS, key=lambda d: -cosine(qvec, DOC_VECS[d]))[:2]
    print("  keyword (BM25): ", ", ".join(keyword))
    print("  semantic:       ", ", ".join(f"{d} {cosine(qvec, DOC_VECS[d]):.2f}" for d in semantic))
Output
Query: "ECONNREFUSED 5432"
  keyword (BM25):  pg-refused 1.85, redis-refused 0.68
  semantic:        redis-refused 0.85, pg-refused 0.82
Query: "app cannot reach the db"
  keyword (BM25):  pool-exhausted 1.24
  semantic:        pool-exhausted 1.00, pg-refused 0.97
  • Exact token query: BM25 ranks pg-refused first because the rare token 5432 appears only in that document, while the vectors confuse the two refused-connection runbooks.
  • Descriptive query: BM25 finds pool-exhausted only through the incidental word "cannot" and misses pg-refused entirely, whereas semantic search retrieves both relevant runbooks.
  • Complementary errors: Each method fails on the query that the other method handles well, which motivates combining them.

Combining Both Methods

In production, lexical search vs semantic search is rarely an either-or choice. Hybrid search runs both retrievers and merges their rankings, and a reranking model, described in reranking in RAG, can then reorder the merged candidates. The nearest-neighbour step inside semantic search is covered in vector search.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. In the Python example, why does semantic search rank redis-refused above pg-refused for "ECONNREFUSED 5432"?

Frequently Asked Questions

What is the difference between keyword and semantic search?

Keyword search matches the exact terms of a query against an inverted index and scores them with a formula such as BM25. Semantic search converts the query and documents into embedding vectors and ranks documents by vector similarity, so it matches meaning rather than wording.

Is BM25 a keyword search algorithm?

Yes. BM25 is a ranking function for lexical search that scores documents by term frequency, the rarity of each term across the collection and document length. It is the default scoring method in many full-text search engines.

Is semantic search always more accurate than keyword search?

No. Semantic search performs better on paraphrased and descriptive queries, but keyword search is usually more accurate for exact identifiers such as error codes, function names and ticket numbers. The better choice depends on the queries users actually type.

Can keyword search and semantic search be combined?

Yes. Hybrid search runs both methods on the same query and merges their rankings, commonly with reciprocal rank fusion. This covers both exact terms and paraphrased intent.