AI visibility guide
Generative Engine Optimization: A Practitioner's Framework
By Jeremy Osborn, Founder and Principal Strategist, Visibility Partners
Published · Last reviewed · ~14 min read
1. What GEO actually is
Generative engine optimization is the practice of making a source more likely to be retrieved, quoted and attributed inside a generated answer, rather than ranked in a list of links.
The term comes from a specific paper: Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, “GEO: Generative Engine Optimization”, published at ACM SIGKDD 2024 (KDD ’24, pp. 5–16, DOI 10.1145/3637528.3671900). Almost every “GEO” claim in circulation traces back to that one paper, usually to a single number from it. It is worth understanding what the number actually measured.
The distinction that organises everything else
Three different outcomes get collapsed into “AI visibility”, and they behave differently.
Three outcomes, three different causes
Citation
Your URL appears as a source under the answer
How it is won
Retrieval — being in the candidate set and surviving reranking
Mention
Your brand is named in the answer text, linked or not
How it is won
The model's parametric knowledge plus retrieved context
Recommendation
Your brand is named as the answer to a choice question
How it is won
Corroboration across independent sources
On Gemini, the overlap between brands mentioned and domains cited can be as low as 30% (Semrush 2026 AI Visibility Index, 126M US prompts, January–April 2026).
Semrush’s 2026 AI Visibility Index (126 million US prompts, January–April 2026) found that on Gemini, the overlap between the brands mentioned and the domains cited can be as low as 30%. Being talked about and being cited are separate results with separate causes. A measurement programme that reports one number for “AI visibility” is hiding the thing you most need to know.
2. What has actually been measured
2.1 The founding experiment, in full
The GEO paper built GEO-BENCH: 10,000 queries (8,000 train / 1,000 validation / 1,000 test) drawn from nine sources including MS Marco, ORCAS-1, Natural Questions, ELI5 and Perplexity.ai Discover, spanning 25 domains, with a query-intent mix of 80% informational, 10% transactional, 10% navigational.
The test harness was a simulated generative engine: take the top 5 Google results for a query, then generate an answer with gpt-3.5-turbo, sampling five responses per query at temperature 0.7. Visibility was scored two ways — Position-Adjusted Word Count, an objective measure of how much of the answer is attributable to a source, weighted by where in the answer it appears; and Subjective Impression, seven sub-metrics judged by GPT-3.5.
Change in Position-Adjusted Word Count vs. no optimization
GEO-BENCH, 1,000 test queries, simulated engine on gpt-3.5-turbo (Aggarwal et al., KDD ’24). Keyword stuffing is the only method that reduced objective visibility.
Source: Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan & Deshpande (2024), GEO: Generative Engine Optimization, KDD ’24, pp. 5–16.
| Method | Position-Adjusted Word Count | Change vs baseline | Subjective Impression | Change vs baseline |
|---|---|---|---|---|
| No optimization (baseline) | 19.3 | — | 19.3 | — |
| Keyword Stuffing | 17.7 | −8.3% | 20.2 | +4.7% |
| Unique Words | 20.5 | +6.2% | 20.4 | +5.7% |
| Authoritative | 21.3 | +10.4% | 22.9 | +18.7% |
| Easy-to-Understand | 22.4 | +16.1% | 20.5 | +6.2% |
| Technical Terms | 22.7 | +17.6% | 21.4 | +10.9% |
| Cite Sources | 24.5 | +26.9% | 21.9 | +13.5% |
| Fluency Optimization | 24.6 | +27.5% | 21.9 | +13.5% |
| Statistics Addition | 25.4 | +31.6% | 23.7 | +22.8% |
| Quotation Addition | 27.3 | +41.5% | 24.7 | +28.0% |
Three things in this table deserve more attention than they get.
The headline “up to 40%” is Quotation Addition, on one metric, in a simulated engine. It is not a traffic figure, a ranking figure, or a real-citation figure. It measures how much of a generated answer was attributable to a source that was already inside the model’s context window. That is a real finding about answer composition. It is not a finding about discoverability.
Keyword stuffing made things worse. It is the only method that reduced objective visibility, from 19.3 to 17.7. The paper’s framing is that tactics effective in search engines may not transfer to generative ones. This is the clearest evidence available that GEO is not SEO with a new label.
The gains concentrate on sources that rank badly. When all five sources were optimised simultaneously, Cite Sources lifted the rank-5 website’s visibility by 115.1% while decreasing the rank-1 website’s visibility by 30.3%. Statistics Addition showed the same shape: +97.9% for rank 5, −20.6% for rank 1. If you are already the dominant source, these tactics work against you. If you are the fifth source in the context window, they are how you get quoted.
The paper also tested 200 samples against live Perplexity.ai. There, Statistics Addition produced the best result — a 37.2% improvement on Subjective Impression. Notably, Cite Sources reversed sign on Perplexity, falling from 24.7 to 19.0. The most SEO-intuitive tactic in the set did not transfer cleanly to the one commercial engine tested.
2.2 The reversal nobody quotes
In 2026 a second team built a more realistic test. Kim, Jeong, Kim, Lee and Lee, “SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization” (arXiv 2602.12187, accepted at KDD 2026), evaluated the full pipeline rather than a fixed context window: retrieval across 171,003 documents and 2,700 queries in nine domains, then reranking, then generation.
Their finding is the single most important correction to practitioner GEO advice:
- Applying the Aggarwal-style body-text rewrites degraded retrieval, with Hit Rate falling roughly 9%.
- Structural information, including schema markup, improved Hit Rate by roughly 22% and gained 2.72 places in average rank.
- Combining both produced +15% at retrieval but −25% Hit Rate at reranking.
The mechanism is straightforward once stated. The 2024 experiment assumed the document was already retrieved. The rewrites optimise for being quoted once you are in the room. Run them in a full pipeline and they can cost you the retrieval step that gets you into the room at all.
2.3 The survey verdict
Olivier Martinez, “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)” (arXiv 2607.14035, 15 July 2026), reviewed 45 studies. It is a single-author preprint and not peer-reviewed, and its conclusion is the honest summary of this field.
The GEO paper’s widely cited gains are, in the survey’s phrasing, “conditional on a source already being present in a fixed context” — and establish neither organic discoverability nor durable traffic effects. Across the reviewed literature, topical relevance and context position are the most reproducible levers; generic heuristics transfer poorly between engines; competitive adoption erodes individual gains; and no reviewed technique demonstrates a stable, longitudinal, cross-platform causal effect on discoverability.
We think that is right, and we would rather say so than sell a certainty that does not exist.
3. Three things the evidence does not support
This section exists because the alternative — repeating claims that have been tested and failed — is how this category loses credibility.
3.1 Schema markup does not measurably increase AI citation
Ahrefs ran the only controlled test we could find: a difference-in-differences study of 1,885 pages that added JSON-LD schema between August 2025 and March 2026, matched against 4,000 control pages, measured 30 days before and after (“We Tracked 1,885 Pages Adding Schema”, 11 May 2026).
| Engine | Effect on citations | Significance |
|---|---|---|
| Google AI Overviews | −4.6% | p ≈ 0.0004 — significant, and negative |
| Google AI Mode | +2.4% | indistinguishable from zero |
| ChatGPT | +2.2% | indistinguishable from zero |
Google’s own documentation agrees. From “AI features and your website” (Search Central, updated 10 December 2025): there is “no special schema.org structured data that you need to add,” and “no additional technical requirements” to appear in AI Overviews or AI Mode.
The study’s authors state its limits, and they matter: the sample was pages already heavily cited by AI, so it says nothing about previously uncited pages; it covered JSON-LD only; the 30-day window may miss slower effects; and schema types were pooled, which would mask any type-specific effect. The −4.6% decline is described as small and unexplained.
The one vendor claim pointing the other way is Microsoft’s Fabrice Canel, speaking at SMX Munich in March 2025, who said schema helps Microsoft’s LLMs understand content. That is a conference remark, about comprehension rather than citation, about one engine, and eighteen months old.
3.2 llms.txt is not read by AI search engines
Ahrefs analysed 137,210 domains via server-log and bot analytics in May 2026 (“We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read”, 15 June 2026). About 38,000 had an llms.txt file. 97% of those files received zero traffic in the month. Of the requests that did occur, only 19.5% came from named AI tools. No AI bot ever requested a file that did not exist — meaning no engine goes looking for it. The authors note every figure is a ceiling on actual consumption.
On the record, from Google: John Mueller, June 2025 — “no AI system currently uses llms.txt.” Gary Illyes at Search Central Deep Dive APAC, July 2025, reported as saying Google does not support it and is not planning to. Google’s written documentation says you do not need to create “machine readable files, AI text files, or markup” to appear in AI features.
OpenAI, Anthropic, Perplexity and Microsoft have said nothing either way. All of them publish an llms.txt of their own, which proves nothing about whether they read anyone else’s.
3.3 Length is not a lever
Ahrefs analysed 174,048 pages with valid data, filtered from 1,677,876 URLs cited across 560,346 AI Overviews (“Short vs. Long Content in AI Overviews”, 3 December 2025). The correlation between word count and citation was Spearman 0.04 — effectively zero. 53.4% of cited pages were under 1,000 words; 16.6% were under 350.
The GEO paper points the same way from a different angle: its winning tactics were about the density and type of evidence in a passage, not its length.
4. What the evidence does support
4.1 Corroboration beats links
Ahrefs correlated AI Overview brand visibility against ten factors across 75,000 brands (“An Analysis of AI Overview Brand Visibility Factors”, 26 May 2025).
Correlation with AI Overview brand visibility
Spearman correlation across 75,000 brands (Ahrefs, 26 May 2025). Correlation is not causation, and every factor shown is a moderate-to-weak relationship.
Branded mentions correlate roughly three times more strongly than backlinks and twice as strongly as Domain Rating. The authors are explicit that this is correlation, not causation, and that all factors show moderate-to-weak relationships. We are repeating that caveat rather than burying it.
But note what a mention is: an independent source saying your brand exists and does a thing. That is corroboration — and corroboration is exactly what a retrieval system needs in order to treat a claim as reliable rather than promotional.
The counter-evidence matters too. Seer Interactive found the correlation between coverage on top news sites and ChatGPT brand mentions to be approximately 0.07 — almost no relationship (3 March 2025; sample size not stated). Their conclusion: being featured in a major publication does not guarantee a brand appears. That study used pre-retrieval training data, so it speaks to what a model recalls rather than what a search layer retrieves — but it is a useful corrective to “get a Forbes placement” as a strategy.
4.2 Being a source of evidence, not a claimant
The one thing the 2024 paper, the 2026 pipeline study and the correlational work all agree on is that quotable, attributable, specific material outperforms assertion. Quotation Addition was the strongest single tactic (+41.5%). Statistics Addition was strongest on live Perplexity (+37.2%). Keyword stuffing was the only tactic that actively hurt.
That is a content instruction, not a technical one: publish material that another party would quote, with figures a model can lift and attribute.
4.3 Position in the candidate set still matters — differently
The rank-position finding in section 2.1 is the most strategically useful result in the literature and the least discussed. GEO tactics deliver their largest gains to sources that are present but not dominant. For a challenger brand already appearing in retrieved context but rarely quoted, the measured upside is large. For a category leader already quoted first, the same tactics can reduce share.
5. How the pipeline actually works, and where each lever acts
Generative answers are produced in stages. Confusing them is the most common analytical error in this discipline.
Where each lever acts in the answer pipeline
Optimizing stage 3 at the expense of stage 1 is the failure mode SAGEO Arena measured: body-text rewrites lifted quotability but cost roughly 9% of Hit Rate at retrieval.
- 1
Retrieval
The query is fanned out into sub-queries and a candidate set is fetched.
- Crawlability for search-layer bots
- Indexation
- Sub-query topical coverage
- Structural clarity
- 2
Reranking
Candidates are scored for relevance to the specific sub-query.
- Passage-level specificity
- Answers placed near their headings
- 3
Generation & attribution
The model composes the answer and attaches citations.
- Quotability
- Statistics with sources
- Clear claim attribution
Stage 1 — Retrieval. A query is expanded into multiple sub-queries and a candidate set is fetched. OpenAI confirms this directly: ChatGPT “typically rewrites your query into one or more targeted queries” (OpenAI Help Center, ChatGPT search). Google calls the same mechanism query fan-out and says it lets Search “display a wider and more diverse set of helpful links” than a standard result page.
Stage 2 — Reranking. Candidates are scored for relevance to the specific sub-query. The levers are passage-level specificity and direct answers positioned near their headings.
Stage 3 — Generation and attribution. The model composes an answer and attaches citations. The levers are quotability, statistics, and clear attribution of claims — the GEO paper’s territory.
The SAGEO Arena result is what happens when you optimise stage 3 at the expense of stage 1. Work the stages in order.
What you can and cannot control at the crawler level
Getting this wrong is common and consequential. From vendor documentation:
| Crawler | robots.txt token | What it governs |
|---|---|---|
| OAI-SearchBot | OAI-SearchBot | Appearance in ChatGPT search results |
| GPTBot | GPTBot | Content that may be used in model training |
| ChatGPT-User | ChatGPT-User | User-initiated fetches in ChatGPT |
| ClaudeBot | ClaudeBot | Content that may contribute to training |
| Claude-SearchBot | Claude-SearchBot | Improving search result quality |
| PerplexityBot | PerplexityBot | Surfacing and linking sites in Perplexity results |
| Perplexity-User | Perplexity-User | User-initiated fetches |
| Google-Extended | Google-Extended | Gemini training and grounding — not AI Overviews |
| CCBot | CCBot | Common Crawl's open repository |
Three corrections worth internalising:
- Blocking GPTBot does not remove you from ChatGPT search. Those are separate opt-outs. OpenAI states that sites opted out of OAI-SearchBot “will not be shown in ChatGPT search answers, though can still appear as navigational links.”
- Google-Extended has no effect on AI Overviews. Google states it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal.” AI Overviews are fetched by Googlebot; the only controls are the snippet controls (nosnippet, data-nosnippet, max-snippet, noindex) — which also degrade normal search appearance.
- Perplexity-User generally ignores robots.txt, because the fetch is user-initiated. Perplexity documents this.
Neither OpenAI nor Perplexity publishes its source-selection criteria. OpenAI says only that ChatGPT “ranks search results using multiple factors” and that “placement is not guaranteed.” Anyone claiming to know the citation algorithm is inferring.
6. Measurement, done honestly
The volatility problem
Ahrefs tracked 43,000+ keywords with a minimum of 16 observations each over a month (“AI Overviews Change Every 2 Days”, 11 November 2025):
- An AI Overview has a 70% chance of changing between consecutive observations.
- Average persistence before the content alters: 2.15 days.
- 45.5% of cited sources are entirely new between consecutive responses.
- But cosine similarity between consecutive versions is 0.95 — the wording and sources rotate; the answer does not.
The authors note their observations were not continuous, so the real change rate is higher.
Three consequences. A single-run check of an AI answer is close to worthless as evidence. Visibility must be reported as a rate across N observations in a defined window, never as presence on a given day. And because the substance is stable while the sources rotate, the durable goal is to become a source of the claim the answer already makes consistently — not to win one citation slot.
The attribution problem
Search Console folds AI Overview and AI Mode data into ordinary Search totals with no filter to isolate them. Any AI visibility measurement has to come from external repeated sampling.
The accuracy problem
The Tow Center at Columbia Journalism Review tested 1,600 queries across eight engines and 20 publishers (“AI Search Has a Citation Problem”, 6 March 2025). Incorrect citation rates: Perplexity 37%, Copilot ~60%, ChatGPT Search 67%, Grok-3 94%, Gemini ~95%. More than half of Gemini and Grok-3 responses cited fabricated or broken URLs. Paid tiers performed worse, because they answered definitively instead of declining.
Separately, an academic audit of AI Overviews (Xu, Iqbal & Montgomery, Washington University in St. Louis, arXiv 2605.14021 — 55,393 queries, 98,020 atomic claims, March–April 2026) found 11.0% of claims unsupported by the pages cited, with omission the dominant failure mode.
Brands are cited for content they did not write, and their content is attributed elsewhere. Any monitoring programme needs to check attribution correctness, not just presence.
What we measure, and why
A defensible measurement design has five properties:
- A locked prompt universe — built once from buyer research, then re-run unchanged. Changing prompts between cycles makes every comparison meaningless.
- Citation and recommendation scored separately — they have different causes and different fixes.
- Repeated sampling within each cycle — because of the 70% change rate, a single observation is noise.
- Presence rate tracked separately from citation rate — platform-wide shifts in how often an engine answers at all will otherwise be misread as your performance changing.
- Attribution accuracy checked — given the error rates above.
7. A 90-day operating model
Days 1–30
Retrieval and baseline
Confirm crawlability for the search-layer bots specifically, not just the training ones. Verify indexation. Build and lock the prompt universe. Take a baseline across engines with repeated sampling. Fix structural clarity: direct answers near their headings, textual content rather than content locked in images, internal links that let a crawler reach the pages that answer sub-queries.
Days 31–60
Evidence
This is where the measured levers live. Publish original data. Make claims quotable and attributable. Add statistics with sources. Cover the sub-queries a fan-out would generate, not just the head term. Resist the temptation to inflate length — the correlation is 0.04.
Days 61–90
Corroboration
Pursue independent mentions rather than links as the primary goal, given the 0.664 versus 0.218 correlation gap — while remembering Seer's 0.07 null result for news coverage specifically. Aim for mentions in the places that already get cited in your category, which the baseline in days 1–30 will have identified. Re-run the locked prompt universe and report the change.
Six to twelve months is the realistic horizon for a category position to move, because that is how long engines take to re-crawl, re-evaluate and re-rank the sources they cite.
8. What nobody knows
An honest framework has to end here.
- No published technique has demonstrated a stable, longitudinal, cross-platform causal effect on organic discoverability. Correlational evidence and single-engine experiments are what exist.
- No engine publishes its selection criteria. OpenAI and Perplexity both decline to.
- We could not locate a peer-reviewed study measuring citation overlap between AI answers and Google organic results. Vendor blogs address it; academic work has not.
- The literature is young. The founding paper is from 2024 and its central result has already been substantially qualified by 2026 work.
Where a practitioner tells you otherwise, ask for the sample size and the date.
Sources
Peer-reviewed and academic
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative Engine Optimization. KDD '24, pp. 5–16. DOI 10.1145/3637528.3671900 · arXiv:2311.09735
- Kim, S., Jeong, W., Kim, S., Lee, S., & Lee, D. (2026). SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization. Accepted, KDD 2026. arXiv:2602.12187
- Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating Verifiability in Generative Search Engines. Findings of EMNLP 2023.
- Xu, H., Iqbal, U., & Montgomery, J. M. (2026). Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact. Washington University in St. Louis. arXiv:2605.14021
- Martinez, O. (2026). Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026). Preprint, not peer-reviewed. arXiv:2607.14035
- Pfrommer, S., Bai, Y., Gautam, T., & Sojoudi, S. (2024). Ranking Manipulation for Conversational Search Engines. EMNLP 2024.
Vendor documentation
- Google Search Central. AI features and your website. Updated 10 December 2025.
- Google Search Central. Google common crawlers.
- OpenAI. Bots.
- OpenAI Help Center. ChatGPT search.
- Anthropic. Does Anthropic crawl data from the web?
- Perplexity. Bots.
- Howard, J. The /llms.txt file, v2. Published 3 September 2024, modified 10 August 2026.
Industry measurement
- Ahrefs (Linehan & Guan). We Tracked 1,885 Pages Adding Schema. 11 May 2026.
- Ahrefs (Linehan & Guan). We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read. 15 June 2026.
- Ahrefs (Linehan & Guan). An Analysis of AI Overview Brand Visibility Factors. 26 May 2025.
- Ahrefs (Linehan & Guan). AI Overviews Change Every 2 Days (But Never Change Their Mind). 11 November 2025.
- Ahrefs (Gavoyannis & Guan). Short vs. Long Content in AI Overviews. 3 December 2025.
- Semrush. 2026 AI Visibility Index. 26 June 2026. 126M prompts.
- Seer Interactive. Does Being Mentioned on Top News Sites Impact AI Answer Mentions? 3 March 2025. Sample size not stated.
- Tow Center / Columbia Journalism Review (Jaźwińska & Chandrasekar). AI Search Has a Citation Problem. 6 March 2025.
- Search Engine Land. Microsoft Bing/Copilot use schema for its LLMs. 20 March 2025. Reported conference remark.
- Search Engine Roundtable. Google Says No AI System Currently Uses LLMs.txt. 17 June 2025.