What is GEO (Generative Engine Optimization)
Generative engine optimization (GEO), also called AI search optimization, is the practice of structuring and publishing web content so that AI answer engines retrieve it, quote it accurately and cite it as a source. For a developer documentation site, it focuses on clear definitions, predictable headings, schema markup and HTML that crawlers can read without running JavaScript.
- Answer engine: A system that generates a written answer from retrieved pages, as explained in how AI search engines work.
- Citation: A link in the generated answer that attributes a statement to a specific source page.
- Extractable passage: A self-contained paragraph or list that remains meaningful when quoted without its surrounding page.
- Schema markup: Structured metadata, usually JSON-LD using the schema.org vocabulary, that describes the page type, title and dates.
- Crawlability: The ability of automated crawlers to fetch the page and read its content from the delivered HTML.
For example, an Orders API documentation site places a one-sentence definition under the H1 of its Rate limits page, so an answer engine can quote that sentence and cite /docs/rate-limits.
Key Characteristics of Generative Engine Optimization
- Passage-level thinking: Answer engines retrieve chunks rather than whole pages, so each section must be understandable independently, as in chunking in RAG.
- Definition-first writing: Each page opens with a precise definition of its subject in one or two sentences.
- Structural clarity: A single H1, descriptive H2 headings and consistent terminology help retrievers identify relevant sections.
- Factual specificity: Concrete values, such as limits, parameters and error codes, give the engine checkable statements to cite.
- Machine-readable metadata: JSON-LD declares the article type, the headline and the modification date.
- Technical accessibility: Server-rendered HTML, permissive robots.txt rules and stable URLs make pages retrievable.
How Generative Engine Optimization Works
- Crawlability: The docs are served as complete HTML. The robots.txt rules allow or block specific AI crawlers on purpose.
- Structure: Every page has one topic, one H1, descriptive H2 sections and a definition sentence at the beginning.
- Specificity: Sections state exact values and procedures, such as the Retry-After header for 429 errors, rather than general descriptions.
- Markup: A JSON-LD block of type TechArticle or FAQPage describes the page, its headline and its dateModified value.
- Discovery files: An XML sitemap lists every page, and an optional llms.txt file summarises the site for language models.
- Monitoring: The team periodically queries answer engines with representative developer questions and records which pages are cited.
The term was introduced in a 2023 research paper that proposed GEO methods and a benchmark for measuring visibility in generative answers.
Example: Auditing a Documentation Page in Python
The program below parses two versions of the same documentation page with the standard HTML parser and checks six GEO signals.
import json
from html.parser import HTMLParser
SERVER_RENDERED = """<html><head><title>Rate limits | Orders API docs</title>
<script type="application/ld+json">{"@type": "TechArticle", "headline": "Rate limits",
"dateModified": "2026-09-01"}</script></head><body>
<h1>Rate limits</h1>
<p>A rate limit is the maximum number of requests a client can send per minute.</p>
<h2>Default limits</h2><p>Each API key can send 600 requests per minute.</p>
<h2>Handling 429 errors</h2><p>Retry after the number of seconds in Retry-After.</p>
</body></html>"""
CLIENT_RENDERED = """<html><head><title>Docs</title></head><body>
<div id="root"></div><script src="/bundle.js"></script></body></html>"""
class PageAudit(HTMLParser):
def __init__(self):
super().__init__()
self.tags, self.paragraphs, self.jsonld, self._tag = [], [], [], None
def handle_starttag(self, tag, attrs):
self.tags.append(tag)
self._tag = "jsonld" if ("type", "application/ld+json") in attrs else tag
def handle_data(self, data):
if self._tag == "p" and data.strip():
self.paragraphs.append(data.strip())
elif self._tag == "jsonld":
self.jsonld.append(json.loads(data))
def audit(html):
page = PageAudit()
page.feed(html)
first = page.paragraphs[0] if page.paragraphs else ""
return {
"one h1": page.tags.count("h1") == 1,
"h2 sections": page.tags.count("h2") >= 2,
"definition first": " is " in first and len(first.split()) <= 40,
"text in HTML": sum(len(p.split()) for p in page.paragraphs) >= 20,
"schema markup": any("@type" in item for item in page.jsonld),
"dateModified": any("dateModified" in item for item in page.jsonld),
}
for name, html in [("server-rendered", SERVER_RENDERED), ("client-rendered", CLIENT_RENDERED)]:
result = audit(html)
missing = [check for check, ok in result.items() if not ok]
print(f"{name}: {len(result) - len(missing)}/{len(result)} checks passed")
if missing:
print(" missing:", ", ".join(missing))server-rendered: 6/6 checks passed
client-rendered: 0/6 checks passed
missing: one h1, h2 sections, definition first, text in HTML, schema markup, dateModified- Server rendering: The server-rendered page exposes its headings, definition, text and JSON-LD in the HTML response, so every check passes.
- Client rendering: The client-rendered page delivers an empty root element, so a crawler that does not execute JavaScript sees no content at all.
- Heuristic checks: These tests approximate structural quality; they cannot measure whether a page is accurate or actually cited.
Applications of Generative Engine Optimization
- API reference documentation: Making endpoint, parameter and error code pages quotable by answer engines.
- Developer tutorials: Structuring step-by-step guides so individual steps can be extracted and cited.
- Troubleshooting guides: Pairing exact error messages with their causes and fixes in dedicated sections.
- Changelogs: Publishing dated, versioned entries so engines can distinguish current behaviour from outdated behaviour.
- Internal knowledge bases: The same structure improves internal semantic search and hybrid search results.
Advantages
- Accurate representation: Clear definitions reduce the risk of an engine paraphrasing the product incorrectly.
- Attribution: Well-structured pages are easier to cite, which directs readers to the authoritative documentation.
- Dual benefit: Work done to optimize content for AI search also improves traditional search results and internal retrieval.
- Maintainability: Consistent page templates simplify documentation reviews and updates.
Limitations
- Opaque ranking: Answer engines do not publish their retrieval and citation criteria, so results cannot be guaranteed.
- Measurement difficulty: Citations vary between runs, engines and phrasings of the same question.
- Evolving conventions: Files such as llms.txt and the handling of AI crawlers are still changing.
- Reduced traffic: A cited answer can satisfy the reader, who then never visits the page.
GEO overlaps with traditional search engine optimization and answer engine optimization; the differences are set out in SEO vs GEO vs AEO.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In the Python example, why does the client-rendered page fail every check?
Frequently Asked Questions
What does GEO stand for in AI search?
GEO stands for generative engine optimization. It is the practice of structuring and publishing content so that AI answer engines can retrieve it, quote it accurately and cite it as a source.
Is GEO different from SEO?
GEO builds on SEO rather than replacing it. SEO aims for a high position in a list of links, while GEO aims for a page's passages to be selected and cited inside a generated answer. Both depend on crawlable, well-structured pages.
What is llms.txt?
llms.txt is a proposed convention for a Markdown file at the root of a website that summarises the site and links to its most useful pages for language models. It is not an official standard, and support among AI systems varies.
Does schema markup help pages get cited by AI search?
Schema markup describes a page's type, title, author and dates in a machine-readable form. It helps systems interpret the page, but no public evidence shows that markup alone causes citations, so clear content remains the priority.
Related Articles
- How AI Search Engines Work (Perplexity, AI Overviews)How AI search engines work: query rewriting, retrieval, passage extraction and cited answer generation, with a Python pipeline over engineering docs.
- SEO vs GEO vs AEOSEO vs GEO vs AEO compared for developer docs: goals, signals, schema markup, crawler rules and llms.txt, with one API docs page optimised all three ways.
- What is Semantic SearchSemantic search explained: how embeddings match meaning instead of exact words, how the pipeline works, where it fails, and a Python runbook search demo.
- Chunking Strategies in RAGChunking strategies in RAG explained: fixed-size, heading-based and semantic chunking, how to choose chunk size and overlap, with a Python comparison.
- Hybrid Search (BM25 + Vector)Hybrid search explained: how BM25 keyword results and vector results are merged with reciprocal rank fusion, with a runnable Python runbook search example.
- Keyword Search vs Semantic SearchKeyword search vs semantic search compared: BM25 term matching versus embedding similarity, a comparison table, when to use each, and a Python BM25 demo.