…
Skip to content
Topics
On this page

How AI Search Engines Work (Perplexity, AI Overviews)

AI search engines are search systems that retrieve documents for a query, read the most relevant passages and generate a single written answer that cites its sources. They apply retrieval-augmented generation (RAG) to a search index, an approach also called generative search, so the response combines the freshness of search with the summarisation abilities of a language model.

  • Query understanding: A language model rewrites or decomposes the question into one or more effective search queries.
  • Retrieval: Keyword, vector or hybrid search returns candidate pages from an index or the live web.
  • Passage extraction: The system selects the specific sentences or chunks that address the question.
  • Grounded generation: A language model composes an answer using only the extracted passages as evidence.
  • Citation: Each statement is linked to the page that supports it, usually with numbered references.
The pipeline behind an AI search engine answerThe question "How do I rotate the signing key?" flows through four stages: the query is rewritten to its key terms, pages are retrieved from a search index, the best passages are extracted, and a language model generates an answer with numbered citations. The answer cites the pages /runbooks/rotate-signing-key and /api/auth. The flow is illustrative.Question: How do I rotate the signing key?1. rewritekey terms2. retrievesearch index3. extracttop passages4. generateLLM + citationsRun keyctl rotate --service orders [1]; requests use HMAC-SHA256 [2][1] /runbooks/rotate-signing-key [2] /api/auth
The pipeline behind an AI search engine answer

For example, an internal AI search tool answers "How do I rotate the signing key?" by retrieving the key rotation runbook and the authentication reference, then citing both pages beside the answer.

Key Characteristics of AI Search Engines

  • Answer-first output: The primary result is a synthesised explanation with supporting source links, which is why these systems are also called answer engines.
  • Multi-query retrieval: Complex questions are decomposed into several sub-queries whose results are consolidated.
  • Attribution: Numbered citations connect individual sentences to specific retrieved documents.
  • Freshness: Answers reflect the current index at query time rather than only the model's training data.
  • Conversational refinement: Follow-up questions reuse the earlier context instead of starting a new search.
  • Probabilistic output: Identical questions can produce differently worded answers from the same sources.

How AI Search Engines Work

  1. Query analysis: The engine interprets the question, extracts key terms and may generate alternative formulations.
  2. Retrieval: A search index returns candidate documents, often dozens, using keyword, semantic search or both.
  3. Reranking: A cross-encoder or similar model reorders the candidates by relevance, as described in reranking in RAG.
  4. Passage extraction: The top documents are divided into passages, and the most relevant ones are selected within a token budget.
  5. Generation: A language model writes the answer from the selected passages and inserts citation markers.
  6. Presentation: The interface shows the answer, the numbered sources and, often, suggested follow-up questions.

Public products such as Perplexity and Google AI Overviews follow this general AI search engine architecture, although their exact ranking signals, models and index sources are not fully documented.

Example: A Toy AI Search Pipeline in Python

The program below retrieves pages by term overlap, extracts the highest-scoring sentence from each page and assembles a cited answer; a template replaces the language model.

Python
import re

PAGES = [
    {"url": "/runbooks/rotate-signing-key",
     "text": "Signing keys are rotated every 90 days. To rotate the key, run "
             "keyctl rotate --service orders. Old keys stay valid for 24 hours."},
    {"url": "/api/auth",
     "text": "Requests are signed with the signing key using HMAC-SHA256. "
             "An expired key returns 401 invalid_signature."},
    {"url": "/runbooks/disk-full",
     "text": "When the disk is full, rotate the logs with logrotate -f. "
             "Then expand the volume."},
]
STOP = {"how", "do", "i", "the", "a", "to", "is", "what", "after", "with"}

def terms(text):
    return {w for w in re.findall(r"[a-z0-9_]+", text.lower()) if w not in STOP}

def search(question, k=2):
    q = terms(question)                                   # 1. query analysis
    scored = [(len(q & terms(p["text"])), p) for p in PAGES]
    print("Retrieval scores:", [(p["url"], s) for s, p in scored])
    hits = [p for s, p in sorted(scored, key=lambda x: -x[0]) if s > 0][:k]  # 2. retrieve
    answer, sources = [], []
    for n, page in enumerate(hits, start=1):
        sentences = re.split(r"(?<=\.)\s+", page["text"])     # 3. extract
        best = max(sentences, key=lambda s: len(q & terms(s)))
        answer.append(f"{best} [{n}]")                    # 4. cite
        sources.append(f"[{n}] {page['url']}")
    return " ".join(answer), sources

answer, sources = search("How do I rotate the signing key?")
print("Answer:", answer)
print("Sources:")
for s in sources:
    print(" ", s)
Output
Retrieval scores: [('/runbooks/rotate-signing-key', 3), ('/api/auth', 2), ('/runbooks/disk-full', 1)]
Answer: To rotate the key, run keyctl rotate --service orders. [1] Requests are signed with the signing key using HMAC-SHA256. [2]
Sources:
  [1] /runbooks/rotate-signing-key
  [2] /api/auth
  • Retrieval cut-off: The disk-full runbook shares only the word "rotate", so it falls outside the top two pages and is never cited.
  • Extraction by overlap: From the authentication page, the sentence with the most query terms is chosen, even though the 401 sentence might be more useful.
  • Traceability: Every sentence in the answer carries a number that maps to one source URL, which a reader can open and verify.

Applications of AI Search Engines

  • Engineering knowledge search: Answering operational questions from runbooks, design documents and postmortems.
  • Codebase question answering: Explaining where and how a function is used, with links to files.
  • Developer documentation: Answering API questions from reference pages with cited sections.
  • Public web research: Summarising several web sources into one cited explanation.
  • Support deflection: Answering customer questions from help-centre articles before a ticket is opened.

Advantages

  • Reduced reading effort: Users receive a synthesised explanation instead of opening many separate pages.
  • Verifiability: Citations allow each statement to be checked against its original source.
  • Current information: Retrieval at query time reflects recently updated documents.
  • Complex questions: Decomposition handles questions that no single page answers completely.

Limitations

  • Misattribution: A citation can point to a page that does not actually support the statement.
  • Hallucination: The model can still add unsupported details, a form of LLM hallucination.
  • Source quality: Outdated or incorrect pages in the index produce outdated or incorrect answers.
  • Latency and cost: Retrieval, reranking and generation take longer and cost more than returning links.
  • Reduced click-through: Readers may open the cited pages less often, which concerns documentation owners and publishers, as discussed in generative engine optimization.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. Which stage of an AI search engine selects the sentences that the answer is built from?

Frequently Asked Questions

How is an AI search engine different from a traditional search engine?

A traditional search engine returns a ranked list of links and snippets for the user to read. An AI search engine also reads the top results and generates a single written answer with citations to the pages it used.

Do AI search engines use RAG?

Yes, in principle. They retrieve documents at query time and pass the relevant passages to a language model as context, which is the retrieval-augmented generation pattern applied to a web-scale or organisation-wide index.

How does Perplexity work?

Perplexity is an AI search engine that runs searches for a question, reads the top-ranked pages and writes an answer with numbered citations to those pages. It follows the general retrieve, extract and generate pipeline, although its exact ranking signals and models are not fully documented.

Can AI search engines give wrong answers?

Yes. The model can misread a source, combine passages incorrectly or rely on an outdated page. Checking the cited sources remains necessary for important decisions.

Can an AI search engine run over private documentation?

Yes. The same pipeline can index an internal knowledge base, codebase or wiki instead of the public web, provided that retrieval respects each user's access permissions.