Resources · Complete guide
Becoming the Answer: How Brands Get Chosen Inside AI Search
The full manuscript by Jeremy Osborn — published free. Fourteen chapters and five appendices on the retrieval pipeline, the five levers of AI visibility, what the evidence actually supports, and how to run the program.
20 chapters · ~26k words · August 2026

About this guide
About the source material
This book began with an academic paper: Answer Engine Optimization: How Agentic AI Reshapes SEO, by Miles Bliey and Keira Chatwin (Stanford GSBGEN 390, July 2026, SSRN). That paper supplies the spine of the argument here — the retrieval pipeline, the Algorithmic Trinity, the NEEATT credibility framework, the distinction between retrieval and validation, and the case that expected rank across a distribution of users has replaced static rank as the relevant visibility metric.
What follows is not a summary of that paper. Every quantitative claim in it was re-verified against primary sources in August 2026, and a meaningful number of them changed. Several were wrong. One widely repeated statistic in the field turns out to be off by a factor of six. Where the paper’s numbers were stale, they have been updated. Where the underlying research has since been challenged — and the central piece of research in this field has been challenged, hard, in a peer-reviewed venue — that challenge is presented in full rather than buried.
That verification pass changed the book. What started as an optimistic manual about a new discipline became something more useful: a manual that tells you which parts of the new discipline are real, which are contested, and which are being sold to you.
How to use this book
The book is in three parts and you can read them out of order.
Part I — The Shift explains what actually changed, how answer engines actually work, and — in Chapter 4 — what the evidence actually supports. If you read only one chapter before approving a budget, read Chapter 4.
Part II — The Five Levers is the implementation half. Five chapters, five levers: entity, retrievability, corroboration, specificity, transactability. Each ends with a playbook you can hand to a team.
Part III — Running It covers measurement, budget, risk, scenarios, and a consolidated 90-day plan.
Every recommendation in Part II and Part III carries an evidence grade:
Grade
Meaning
A
Multiple independent measurements, or a first-party platform specification. Act on it.
B
One good study or consistent practitioner measurement, with a plausible mechanism. Act on it, verify locally.
C
Contested. Reasonable people disagree, or the only evidence is correlational and confounded. Do it if it is cheap.
D
No evidence, or evidence against. Do not fund it.
F
Evidence of harm, or unlawful. Prohibit it.
A plus or minus modifies within a grade — B+ is a strong B, C− a weak C.
The grades are the most valuable thing in the book. This field has a great deal of confident advice and very little confirmed knowledge, and the gap between those two things is where marketing budgets go to die.
A note on sources: where a study was produced by a company selling software in this category, it is labelled as vendor research. That does not make it wrong. It makes it interested, and interested research deserves a heavier thumb on the scale.
PART I — THE SHIFT
Chapter 1: The End of the Click Compact
The deal that held for twenty-seven years
In 1998, PageRank established an arrangement so successful that almost nobody noticed it was an arrangement. Search engines would organize the world’s information. Publishers and brands would produce it. Users would travel between the two through a ranked list of blue links. Nobody signed anything. It simply worked, and an industry worth more than a trillion dollars a year in advertising grew on top of it.
The compact had a single load-bearing assumption: the search engine was a passthrough. Value was created at the click. Everything marketers built — rank tracking, click-through-rate optimization, attribution modeling, the entire apparatus of performance marketing — assumed that a user’s journey began at a search box and ended somewhere else, and that the somewhere else was measurable.
That assumption is now failing, and it is failing in a stranger way than most of the industry admits.
What actually happened
Start with adoption, because the adoption numbers are real and they are large.
ChatGPT approached one billion weekly active users by July 2026, up from 800 million in October 2025 and 900 million in February 2026 — a milestone it reached roughly seven months later than OpenAI had projected, but reached nonetheless. Google’s Gemini app crossed one billion monthly active users in August 2026, making it the fourteenth Google product to do so. Across all generative AI platforms, Similarweb counted 9.5 billion monthly web visits in mid-2026, up 70% year over year, from 655 million unique visitors.
The competitive picture is moving fast. ChatGPT’s share of generative-AI web traffic fell from roughly 76% in June 2025 to about 53% in May 2026, while Gemini rose from under 9% to roughly 27–28% and Claude from about 2% to 9%. Perplexity and Copilot remain in low single digits. Whatever else is true, this is not a market converging on a single gateway. It is fragmenting.
So far, so familiar. Here is the part the conference keynotes leave out.
AI assistants have almost entirely added to search rather than replaced it. Similarweb found that 95% of ChatGPT users also still use Google — a figure unchanged from September 2025 to August 2026. People did not switch. They added a tool.
AI platforms send almost no traffic. Ahrefs, measuring roughly 82,000 websites, found AI referrals averaged 0.25% of total traffic. Microsoft’s own Clarity study put it at “less than 1%.” Scrunch, tracking millions of events across news publishers, found 1.1% of visits carried an AI referrer, against roughly 9% from traditional search and 75% direct. The growth rates are spectacular — Ahrefs measured a 9.7x year-over-year increase, Adobe measured 393% growth in AI-referred retail traffic in Q1 2026 — but they are spectacular growth from a rounding error.
And yet organic traffic is genuinely collapsing. SparkToro, analyzing Similarweb clickstream data, found that 68.01% of US Google searches ended without a click in the first four months of 2026, up from 60.45% in 2024 and 49% in 2019. Pew Research — the only major measurement here with no commercial interest in the answer — observed real browsing behavior across 68,879 searches and found users clicked a result in 8% of visits when an AI summary appeared, versus 15% when it did not. They clicked a link inside the AI summary in 1% of visits. Ahrefs, comparing 300,000 keywords across two years of aggregated Search Console data, measured position-one click-through rate on AI Overview keywords falling 58% against the counterfactual. Publishers report roughly a third less referral traffic year over year.
Put those three findings side by side and the actual shape of the disruption appears:
The traffic loss is happening inside Google’s own results, not because of chatbot referrals. The chatbots are barely sending traffic in either direction. Google’s AI features are absorbing the clicks that used to leave.
This matters enormously for how you allocate. A marketing team that responds to this shift by building a ChatGPT-referral acquisition program has misread the situation. The referrals are not the prize. Being present in the answer — where nobody clicks at all — is the prize, and the fact that it produces almost no measurable traffic is precisely what makes it hard to fund and easy to ignore.
The quality paradox
The traffic that does arrive from AI platforms behaves differently, and here the evidence is genuinely encouraging — with a caveat that most people quoting it omit.
Adobe Digital Insights, working with large-scale retail data, found AI-referred traffic converting 42% better than non-AI traffic in March 2026, with 48% more time on site, 37% higher revenue per visit, and a 12% lower bounce rate. Microsoft’s Clarity study of 1,277 domains found LLM-referred visitors signing up at 1.66% versus 0.15% from search. Ahrefs, on its own site, saw 0.5% of visitors from AI search drive 12.1% of signups. The Washington Post’s chief revenue officer reported AI-platform visitors subscribing at four to five times the rate of search visitors.
The caveat: twelve months earlier, Adobe’s own data showed AI-referred traffic converting at roughly half the rate of everything else. The conversion advantage is recent. It is not a law of nature, it is a snapshot of a moment when the population using AI assistants for purchase research skews toward high-intent, high-income, technically confident users. As adoption broadens, expect the advantage to compress.
And the picture is not uniformly flattering. Ahrefs, across roughly 82,000 sites, found AI visitors bouncing more than search visitors (67.8% versus 63.7%) and viewing fewer pages (4.0 versus 5.2). In news specifically, Scrunch found AI Overview-present searches produced publisher clicks about 20% of the time versus 30% without.
The honest synthesis: AI referral traffic is a small, volatile, currently high-intent slice. The vendor claims of 3x, 5x, and 23x uplift come from small or first-party samples and should be treated as directional. Adobe’s +42% is the number to plan against, and it should be planned against as a figure that may decay.
The Perfect Click, examined
The strategic argument that follows from all this has a name: the Perfect Click. When an AI assistant researches options, compares them, and recommends one, the user who then clicks through is not exploring. The decision is substantially made. That click is confirmatory, and one confirmatory click may be worth more than hundreds of unqualified impressions.
The argument is sound and the conversion data supports it. But it is worth being precise about what it implies, because it is routinely over-read.
It does not imply that AI visibility will replace your traffic. At current volumes it cannot. It implies something subtler and more uncomfortable: that a growing share of purchase decisions is being made in a place you cannot see, cannot measure with existing tools, and cannot buy your way into. The 68% of searches that end without a click did not stop happening. They stopped being observable.
That is the real transition. Not from clicks to no-clicks. From a measurable funnel to an unmeasurable one.
What actually changed, precisely
Strip away the noise and four things changed:
1. The unit of competition moved from the page to the passage. Retrieval systems chunk documents into segments — a few hundred tokens each — and embed and rank them independently. Your 3,000-word guide does not compete as a guide. At the chunk sizes documented in Chapter 2, it competes as roughly eight separate fragments, most of which will never be seen together.
2. The search space expanded. A single user prompt triggers multiple retrieval events against synthetic sub-queries the user never typed and you will never see in a keyword tool. You compete across a distributed set of latent questions.
3. Ranking became comparative. Passages are evaluated against specific competing passages rather than scored in isolation. There is no absolute quality bar. There is only “better than the other thing in the context window.”
4. Validation became a separate gate from retrieval. Being findable is necessary and no longer sufficient. Content that cannot be corroborated against other sources gets discounted during grounding, and answers with low confidence can be suppressed entirely.
Chapter 2 works through the machinery that produces those four properties. Chapter 4 tells you which of the tactics they imply actually survive contact with evidence.
One more thing worth knowing before you plan
Advertising is arriving inside the answer. OpenAI began testing ads in ChatGPT in February 2026, expanding through August, on free and lower-cost tiers, with sponsored content labelled and visually separated. By mid-2026 Similarweb measured roughly 26% of ChatGPT responses containing an ad.
Any strategy that assumes AI answers will remain a purely organic surface is planning against a world that is already ending. The organic and paid layers will separate, as they did in search — but the surface is smaller, the answers are shorter, and the space above the fold in a chat interface is far more contested than a search results page.
Chapter 2: Inside the Answer Engine

You do not need to be an engineer to run an AEO program, but you do need an accurate mental model of what happens between a user’s question and the answer they read. Almost every bad tactic in this field comes from an inaccurate one.
What follows is assembled from three kinds of evidence, and it is worth knowing which is which: platform documentation (Google Search Central, Vertex AI docs, OpenAI’s commerce specs — authoritative), patents (real, but a patent describes what a company thought worth protecting, not necessarily what it shipped), and practitioner reverse-engineering (useful, and clearly labelled as inference).
Step 1: Query understanding
The system converts the user’s prompt into vector representations capturing intent, entities, and task type. It also incorporates context: prior conversation turns, stated preferences, sometimes account history and device state.
The strategic implication arrives immediately and it is large. The same question asked by two different people can trigger different retrieval paths and produce different answers. There is no single results page that everyone sees. Chapter 10 deals with what that does to measurement; for now, hold onto the fact that “our rank” is no longer a coherent concept.
Step 2: The retrieval decision — the gate nobody talks about
Before any of the interesting machinery runs, the system decides whether to search the web at all.
This is documented. Google’s Grounding with Google Search assigns each prompt a prediction score between 0 and 1 estimating whether grounding would help, with a default threshold of 0.3. Below it, the model answers from what it already knows, with no retrieval and no citations.
The observed rates are striking. Profound, analyzing roughly 730,000 US English ChatGPT conversations from late 2025, found only 18% triggered a web search at all. Semrush, using clickstream data through February 2026, found web search enabled on 34.5% of ChatGPT queries — down from 46% in late 2024.
Sit with that. Somewhere between two-thirds and four-fifths of prompts are answered from the model’s parametric memory. No amount of on-page optimization wins a query the system decided not to search for. For those queries, what matters is whether the model already knows who you are — which is a training-data and brand-prevalence question, not a content question. It is the strongest argument in this book for the unglamorous work of Chapter 5.
Step 3: Query fan-out
When the system does retrieve, it typically does not run one search. It decomposes the prompt into multiple synthetic sub-queries representing different facets of intent, and runs them in parallel.
This is officially confirmed. Google Search Central describes query fan-out as the mechanism by which AI features surface “a wider and more diverse set of helpful links.” A Google engineering director put it colloquially on the company blog in March 2026: “AI Mode is basically doing a dozen searches for you in the time it takes to do one.” Google’s patent Search with stateful chat describes generating “one or more synthetic queries” — including alternative suggestions, supplemental queries, rewrites, and drill-down queries — and retrieving against both the original and the synthetics. Vertex’s grounding metadata exposes a webSearchQueries array, so the API itself admits plurality.
What is not documented is how many. The widely circulated figures — eight sub-queries, twelve, hundreds — have no primary source. The only official quantity is that colloquial “a dozen,” in a consumer blog post about visual search. The two most technically credible practitioner analysts in this field, Michael King and Olaf Kopp, both decline to state a number. So should you. The mechanism is real; the arithmetic is invented.
The practical consequence stands regardless: you are competing on questions nobody typed. A prompt like “best CRM for a 12-person agency that does retainer work” fans out into questions about pricing tiers, agency-specific features, retainer billing, small-team onboarding, and integration coverage. You may win four of those and lose on price, and lose the answer.
Step 4: Passage retrieval — why the page is no longer the unit
Each sub-query returns candidate passages, not pages. Documents are chunked into semantically coherent segments, each independently embedded and indexed.
The chunk sizes are documented, and the numbers commonly quoted in SEO circles are too narrow. Google’s Vertex AI Search chunker defaults to 500 tokens with a supported range of 100–500, with an option to include document headings in each chunk explicitly “to prevent context loss in chunk retrieval and ranking,” and retrieval that can return the matched chunk plus up to five adjacent chunks on either side. Google’s Vertex AI RAG Engine defaults to 1,024 tokens with 256 tokens of overlap. Microsoft’s Azure AI Search recommends starting at 512 tokens with 25% overlap.
No vendor discloses what chunk size Google Search itself uses in production. The Vertex figures are a proxy, not a disclosure. The defensible statement is: commonly a few hundred tokens — Google caps its search chunker at 500, Azure recommends starting at 512.
Two things follow. First, a passage that opens “As we discussed above” or “Building on the previous point” carries a dependency that may not survive chunking; retrieved in isolation it is ambiguous, and ambiguity loses. Second — and this is the underappreciated half — the adjacent-chunk and heading-inclusion features mean context is not entirely lost. Structure helps, but the catastrophic-decontextualization story that some AEO decks tell is overstated.
Step 5: Reranking — the comparison you never see
Candidate passages get reranked before they reach the model’s context window, and there is good reason to believe pairwise comparison is involved.
The research is solid. Google Research published Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting in Findings of NAACL 2024, showing that asking a model to compare two documents at a time — a far simpler task than scoring one in isolation or ordering a whole list — lets a 20-billion-parameter open model match GPT-4’s reranking quality. The paper describes aggregating those comparisons via heapsort (O(N log N)) or sliding passes (O(N)). Google then patented the method, with the same author list.
The honest framing matters here. This is an information-retrieval methods paper that Google found valuable enough to patent. It is not a description of AI Mode’s production ranker, and anyone telling you “Google ranks your page by pairwise comparison” is stating an inference as a fact. What you can defensibly say: Google is demonstrably interested in LLM-based pairwise reranking, and the architecture implies a competitive rather than absolute standard.
The strategic reading of that: your passage does not need to be good. It needs to beat the specific passage it is compared against. A verifiable number, a named source, a concrete example — these give a reranker something to prefer. A paraphrase of common knowledge gives it nothing. Practitioners call the difference information gain, and it is the single most transferable idea in the discipline.
Step 6: Grounding and validation
Surviving passages are checked against the answer being drafted. This is well documented in Google’s Check Grounding API, and the specifics are instructive:
A support score from 0 to 1 indicating how grounded an answer candidate is in the provided facts — roughly, the fraction of claims found supported.
A citation threshold, default 0.6, controlling how confident a citation must be to appear.
Claim-to-chunk mapping at roughly sentence granularity, with optional per-claim scoring.
Google’s granted patent on generative summaries goes further: confidence for each portion of the summary is derived from the model’s own confidence, the trustworthiness of supporting documents, and how many documents verify that portion. Low-confidence summaries can be suppressed entirely.
Read that last mechanism again, because it is the technical basis for the whole second half of this book. Corroboration count is an input to whether your claim survives. A fact asserted only on your own website is a single-source claim. The same fact, appearing consistently on your site, in a trade publication, in a directory listing, and in a Wikidata statement, is a multi-source claim. The architecture prefers the second, and it prefers it mechanically, not as a matter of editorial taste.
Step 7: Generation and attribution
Validated passages go into the context window with source identifiers embedded, letting the system link specific spans of the answer back to specific documents — sometimes as page links, sometimes as anchor links to a specific passage.
Output is probabilistic. The same prompt run twice produces different answers, and the variance is larger than most people assume. A University of St. Gallen study found day-to-day source overlap for identical prompts at Jaccard 0.34–0.42, and brand-mention overlap at 0.45–0.59 — with the instability persisting within 24-hour windows, ruling out news churn as the explanation. Chapter 10 explains what that does to your reporting. The short version: one measurement is not a measurement.
A correction worth making out loud
Two patents circulate in AEO presentations as evidence of Google’s AI search architecture. One of them is not Google’s.
US20240346256A1, “Response generation using a retrieval augmented AI model,” is assigned to Microsoft Technology Licensing, not Google. It is also unremarkable: it describes plain single-query vector RAG with cosine similarity and a top-K threshold, containing no query decomposition or fan-out at all. If you see it cited as Google evidence for fan-out, the deck is wrong twice.
The patent you actually want is US20240289407A1, “Search with stateful chat.” The mistake is a small one, but it is a useful diagnostic: it tends to appear in material that was assembled from other people’s summaries rather than from sources.
What this machinery implies
Mechanism
What it means for you
Retrieval threshold (default 0.3)
Most prompts never search. Prior brand knowledge matters more than content for those.
Query fan-out
You compete on questions nobody typed and no keyword tool shows.
Passage-level chunking (~500 tokens)
Sections must stand alone. The page is not the unit.
Pairwise reranking
Relative, not absolute. Information gain beats polish.
Grounding with corroboration counting
Third-party agreement is a mechanical input, not a nicety.
Confidence suppression
Ambiguous or contradicted entities can be dropped entirely.
Non-deterministic generation
Single-point measurement is noise.
Chapter 3: The Three Systems You Have to Satisfy
Jason Barnard’s useful contribution to this field is the observation that an answer engine is not one system but three, and that being visible in one of them is worth very little on its own. He calls it the Algorithmic Trinity.
The language model is the synthesis layer. It does not independently verify anything. It reasons over whatever passages arrive in its context window and turns them into readable prose. It is, in the most literal sense, only as good as what it was handed.
The search index is the retrieval layer. It decides what gets handed over. It does not generate answers; it identifies raw material.
The knowledge graph is the validation layer. A structured database of entities, their attributes, and their relationships. It answers a prior question: is this thing real, and is it the thing being discussed?
Your content must be retrievable by the second system, extractable by the first, and your organization must exist as a resolvable entity in the third. Fail any one and the other two cannot save you.
Understandability comes before credibility
Google’s E-E-A-T framework — Experience, Expertise, Authoritativeness, Trustworthiness — has organized thinking about search quality since 2018 and remains useful. But it has a structural gap when the evaluator is a machine.
A machine cannot apply a credibility signal to an entity it cannot identify. If the system cannot resolve who you are, it has nowhere to attach your expertise. Credibility requires a prior step: understandability.
This is the logic behind extending E-E-A-T to NEEATT, adding Notability and Transparency:
Notability — the degree to which you are recognized across independent third-party sources. It functions as a multiplier: the same claim from the same expert is weighted differently depending on whether that expert appears across many credible datasets. Critically, notability is comparative, not absolute. A specialist firm can be more notable within a narrow domain than a conglomerate that is diffusely present across many. This is the structural opening for smaller brands and it is real.
Experience — evidence of actual real-world involvement. Case studies, first-hand accounts, practitioner data, original measurement. This is the signal that distinguishes human knowledge from synthetic summary, and as the web fills with generated content it is becoming more valuable, not less.
Expertise — technical depth. In a pairwise comparison, a passage demonstrating genuine mastery beats a competent overview of the same topic. Depth is a competitive weapon in a way it never quite was under keyword-era SEO.
Authoritativeness — third-party validation. Citations, references, endorsements from independent sources. Note the mechanism from Chapter 2: agreement across multiple documents is an input to grounding confidence. Authoritativeness is not a metaphor here; it is arithmetic.
Trustworthiness — historical reliability. Contradictory claims and inconsistent data erode it, and erosion is expensive to reverse.
Transparency — clarity about ownership, authorship, and intent. If the system cannot parse who owns the content and what entity it represents, it cannot map it to the knowledge graph, and none of the other five signals can be applied at all.
Read in order, the framework says something specific: transparency is a prerequisite, notability is a multiplier, and the middle four are the substance. Most brands invest in the middle four and neglect both ends.
The Entity Home
If the knowledge graph needs to resolve you to a thing, something has to be the canonical statement of what that thing is. Barnard calls this the Entity Home: a single, authoritative, brand-owned property where the system finds definitive facts about your identity, purpose, and activities.
In practice it is usually your homepage or a dedicated About page. But the concept is more demanding than “having a website.” An Entity Home has four properties:
Explicit machine-readable identity. What the organization is, what it does, what category it belongs to, who it serves — stated plainly, in text, not implied by design.
Structured data that maps to knowledge graph entity types. Organization schema at minimum, with sameAs pointing at your verified profiles and identifiers.
Maintenance. Reliably current. A stale Entity Home teaches the graph that your facts are unreliable.
External corroboration. If your homepage says you are an enterprise SaaS company specializing in supply-chain logistics, that same characterization must appear on your LinkedIn page, your Crunchbase listing, your press coverage, your directory entries, and your Wikidata item.
That fourth property is where nearly everyone fails, and the failure mode is boring: a repositioning that reached the website and the sales deck but never reached the eleven third-party profiles created by three former employees over six years. The knowledge graph builds confidence through corroboration. Inconsistency does not merely fail to help — it actively depresses the confidence score, and low confidence triggers the suppression mechanism described in Chapter 2.
Chapter 5 turns this into a work plan.
Two jobs, not one
The Trinity implies that AEO is two distinct disciplines that are usually conflated:
Content optimization makes information citable. It is about passages, structure, specificity, and information gain. It is largely within your control and it is largely a publishing problem.
Brand optimization makes the entity recognizable and verifiable. It is about identity, consistency, third-party corroboration, and prevalence. It is largely outside your direct control and it is a communications and operations problem.
Most teams staff the first and ignore the second, because the first looks like SEO and reports to someone who already exists. The evidence in Chapter 4 suggests the allocation should run the other way.
Chapter 4: What the Evidence Actually Supports
This chapter exists because this discipline has a research problem. It has plenty of research — the volume is genuinely impressive — but a large share of it is produced by companies selling tools in the category, most of it is correlational, and the single most-cited study in the field has since been substantially contradicted in a peer-reviewed venue by a larger one that almost nobody quotes.
If you are going to sign off on budget, you need the scoreboard. Here it is.
The famous study, read properly
Everyone in this field cites the GEO paper: GEO: Generative Engine Optimization, by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, from Princeton, Georgia Tech, IIT Delhi and the Allen Institute, published at ACM SIGKDD 2024. It tested nine content optimization methods across GEO-Bench, a benchmark of 10,000 queries spanning 25 domains.
Here is what it actually found, from Table 1, against a baseline of 19.3 on both metrics:
Three corrections to what you have probably been told:
| Method | Position-adjusted word count | vs. baseline | Subjective impression | vs. baseline |
|---|---|---|---|---|
| Add quotations | 27.2 | +41% | 24.7 | +28% |
| Add statistics | 25.2 | +31% | 23.7 | +23% |
| Optimize fluency | 24.7 | +28% | 21.9 | +13% |
| Cite sources | 24.6 | +27% | 21.9 | +14% |
| Use technical terms | 22.7 | +18% | 21.4 | +11% |
| Make easy to understand | 22.0 | +14% | 20.5 | +6% |
| Authoritative tone | 21.3 | +10% | 22.9 | +19% |
| Unique words | 20.5 | +6% | 20.4 | +6% |
| Keyword stuffing | 17.7 | −8% | 20.2 | +4% |
“Statistics improve visibility by up to 40%” is wrong. Statistics delivered +31%. The best-performing method was quotations, at +41%, and the “up to 40%” figure is the abstract’s rounded headline for the method set as a whole — not a result for statistics. This misattribution appears in the source paper for this book and in a great deal of industry material.
“Fluency optimization improves visibility 15–30%” attaches the wrong number to the wrong method. The 15–30% range is in the paper, but it is the Subjective Impression band for the top-performing methods collectively — not a result for fluency optimization, which scored +28% and +13% on the two metrics. This is the most common way statistics in this field go wrong: a real number, detached from what it measured.
“Keyword stuffing showed negligible improvement” understates it. Keyword stuffing was the worst-performing method tested and it went backwards — −8%. You can say this more strongly than the industry does.
Now the caveats, which are more important than the corrections:
The main experiments did not run on Google, Bing, or any deployed engine. The authors built their own generative engine: retrieve sources, generate a response with GPT-3.5-Turbo. A smaller validation subset ran on Perplexity.
All gains are conditional on already being retrieved into a fixed context of roughly five documents. The study measures how much of the answer your text wins once you are in the room. It says nothing about getting into the room.
The evaluation judge was GPT-3.5 — LLM-as-judge, not human annotation.
It was written against GPT-3.5-era systems in late 2023. It is now the oldest major work in the field.
So the honest reading of the most famous study in AEO is: content rewriting can shift how much of an answer your text wins, in a simulated engine, once you have already been retrieved.
The study that contradicts it
At NeurIPS 2025 — peer-reviewed, Datasets and Benchmarks track — Puerto, Gubri, Green, Oh and Yun published C-SEO Bench, purpose-built to test whether GEO-style methods work. Two tasks (question answering and product recommendation), three domains each, and critically, multiple adoption-rate scenarios modelling what happens when competitors optimize too.
The abstract states the finding plainly: “Most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking.” And: “Traditional SEO strategies, those aiming to improve the ranking of the source in the LLM context, are significantly more effective.”
The adoption-rate result deserves particular attention from anyone building a business case. As more competitors adopt these tactics, individual gains diminish. The game is substantially zero-sum. A tactic that works because few people use it is not a strategy; it is a window.
C-SEO Bench is also a constructed benchmark rather than a deployed engine — the same limitation the GEO paper carries, and it should be applied to both. But it is larger, more recent, more adversarially designed, published in a stronger venue, and it largely fails to replicate the industry’s founding result. Any chapter of any book built on the GEO paper needs this one sitting next to it.
The largest controlled experiment
Vishwakarma, Kumar and Jamidar, publishing at SIGIR 2026, ran 252,000 trials across six models in a controlled two-document RAG testbed, varying exactly one of eighteen factors at a time with anonymized brands.
The factors that dominated, with odds ratios above 10,000:
Topic relevance
Presence of price information
Recent timestamps
Position in the candidate list
Moderate effects: missing specifications, hedged versus confident tone, evidence-backed claims.
And the finding that should reorganize a lot of content roadmaps: formatting and content structure showed minimal effect across all models.
That is an uncomfortable result for an industry that has spent two years selling bullet-point restructuring, question-formatted H2s, and “chunk-friendly” rewrites. Two caveats. The study is affiliated with Sprinklr, so it is vendor-adjacent. And like the two studies above it, it runs on a synthetic testbed rather than a deployed engine — the odds ratios describe what moves a controlled reranker, not a measured lift in Google. With those attached, the design is unusually rigorous and it appeared at SIGIR.
The largest audit of real AI Overviews
Xu, Iqbal and Montgomery at Washington University in St. Louis audited 55,393 trending queries across 40 days in spring 2026, yielding 7,583 AI Overviews, and verified 98,020 atomic claims against their cited sources.
Finding
Figure
AI Overview activation rate overall
13.7%
Activation on question-form queries
64.7% (versus 9.5% on non-question queries — 6.8x)
Category range
3.5% (beauty/fashion) to 46.1% (hobbies/leisure)
Suppressed categories
Politics 7.5%, law/government 9.6%
Cited domains absent from the co-displayed first page
29.8%
Claims inconsistent with their cited sources
11.0% (4.1% contradicted, 7.0% not present in source)
UGC share of AI Overview citations
14.2% (versus 41.4% of first-page results)
Two things to take from this. First, roughly 70% of AI Overview citations come from domains already on page one — traditional ranking is a strong but incomplete predictor, and the incompleteness is the opportunity. Second, the 6.8x activation multiplier on question-form queries is one of the few well-evidenced, actionable content facts in the entire literature — though note carefully that it concerns the user’s query form, not your page’s headings.
The 11% claim-inconsistency rate is a separate matter entirely, and Chapter 12 deals with it.
The measurement bomb
Schulte, Bleeker and Kaufmann at the University of St. Gallen tested something nobody selling dashboards wanted tested: run the same prompt repeatedly and see how much the answer changes.
Day-to-day source overlap for identical prompts: Jaccard 0.34–0.42. Brand-mention overlap: 0.45–0.59. The instability persists within 24-hour windows, ruling out news cycles as the cause. It is model stochasticity.
Their recommendation: at least seven runs per prompt per day for brand tracking to get standard error below 0.10, at least eight for source coverage, and rolling windows of 14 to 21 days.
Their conclusion: “single observations of AI visibility are misleading.”
Most commercial AI-visibility tools sample far less than this, and almost none publish confidence intervals. Which means: any before-and-after case study claiming a 30% visibility improvement, without disclosed repeated sampling, is inside the noise floor. Including, potentially, yours. Chapter 10 tells you how to avoid producing one.
The schema question, settled as well as it can be
Structured data is the most confidently asserted tactic in this field, so it is worth knowing that the best causal study available finds essentially nothing.
Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against 4,000 matched control pages, using matched difference-in-differences with event-study and symmetrical-window analysis over 30-day pre/post windows.
| Platform | Effect of adding schema |
|---|---|
| Google AI Overviews | −4.6% (small, statistically significant, unexplained) |
| Google AI Mode | +2.4% (indistinguishable from zero) |
| ChatGPT | +2.2% (indistinguishable from zero) |
The authors’ explanation for the widely cited correlation between schema and AI citations: confounding. Well-maintained sites do structured data and everything else.
A separate retrieval-mechanics test by searchVIU is even more direct. A test page carried eight product prices distributed across visible HTML, JavaScript-rendered content, JSON-LD only, hidden Microdata, visible Microdata and RDFa. No system found prices that existed only in JSON-LD. Gemini scored highest at 50% — the only system rendering JavaScript during live fetch. ChatGPT managed 37.5%, reading visible HTML only. Claude scored zero.
And Google’s own documentation says, verbatim: “You don’t need to create new machine readable files, AI text files, or markup to appear in these features. There’s also no special schema.org structured data that you need to add.”
This does not mean abandon schema. It means stop funding it as a citation lever. Schema retains genuine value for entity disambiguation (Chapter 5), for Bing and Copilot, for classic rich results, and for commerce surfaces where the feed specification is explicit. Those are real jobs. “Getting cited in AI answers” is not one of them.
The technical finding that is not contested at all
Amid all this uncertainty, one technical result is measured repeatedly, by independent parties, with no disagreement.
Most dedicated AI crawlers do not execute JavaScript.
Vercel, measuring roughly 1.3 billion monthly AI crawler requests, found no major AI crawler executing JavaScript. They fetch JS files — 11.5% of ChatGPT’s requests, 23.8% of Claude’s — and never run them. searchVIU analyzed 23 major AI crawlers in November 2025 and found 69% cannot execute JavaScript — which also means 31% can, with Googlebot, Bingbot and Gemini’s live fetch the notable exceptions. Note the dates: Vercel’s measurement is December 2024 and searchVIU’s is November 2025, in the fastest-moving area this book covers. Re-test rather than assuming.
Vercel found something else worth acting on: ChatGPT spends 34.8% of its fetches on 404s and another 14.4% on redirects. Claude burns 34.2% on 404s. Googlebot’s figures are 8.2% and 1.5%. Roughly a third of the AI crawl budget spent on your site is being wasted on dead URLs.
This is the highest-confidence technical recommendation in the book: server-side render or statically generate anything you want cited, and clean up your 404s. It is unglamorous, it is cheap, and unlike almost everything else in this chapter, nobody disputes it.
The cargo cult
llms.txt. John Mueller of Google, June 2025: “FWIW no AI system currently uses llms.txt.” A server-log study across roughly 900 domains from September 2025 to April 2026 recorded 1,227 total llms.txt requests — fewer than seven per day across all sites — of which 64.7% came from a single data-broker crawler and 31.9% from Chrome browsers, meaning humans. Zero of 1,227 requests came from a frontier AI lab’s crawler.
The one genuine signal: Chrome added an llms.txt Lighthouse audit in May 2026, filed deliberately under “agentic browsing” rather than SEO. That is where the file may eventually matter — agent tooling, not search visibility.
Publishing one costs an hour and harms nothing. Do not let anyone put it on a roadmap as a visibility driver, and do not pay an agency for it.
Blocking Google-Extended to control AI Overviews. It does not work, because it is not that kind of control. Google-Extended is not a crawler; it is a robots.txt token governing whether content may be used to train and ground Gemini Apps and the Vertex API. Google’s documentation states it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal.” AI Overviews and AI Mode are features within Google Search, so the token does not touch them. The only lever is the Search Console opt-out toggle, and Chapter 10 explains why you should not pull it.
The scoreboard
The uncomfortable through-line
Read together, the strongest studies point the same direction, and it is not the direction the industry sells:
Getting into the retrieval set dominates everything that happens afterward.
Keyword stuffing
Actively negative
D
| Tactic | Evidence | Grade |
|---|---|---|
| Server-side rendering / content in initial HTML | Multiple independent measurements, no dissent | A |
| Fixing 404s and redirect chains | Vercel crawl-waste data; mechanically obvious | A |
| Crawler access control by correct user agent | First-party platform documentation | A |
| Product feed completeness for commerce surfaces | Google and OpenAI first-party specs | A |
| Entity establishment (Wikidata, sameAs, consistency) | Mechanism documented; strong correlational support | A− |
| Earned third-party mentions | Strongest correlate in the largest correlational study | B+ |
| Organizing coverage around real question-form queries | Xu et al., 6.8x activation multiplier | B |
| Original data, statistics, quotations, named sources | GEO paper; partially challenged by C-SEO Bench | B− |
| Passage self-containment | Mechanism documented; direct effect unmeasured | B− |
| YouTube presence | Strongest single correlate (r ≈ 0.74); causality unestablished | B− |
| Structural rewriting for “chunkability” | Sprinklr: negligible. C-SEO Bench: sometimes harmful | C |
| JSON-LD schema as a citation lever | Best causal study finds ~0 or slightly negative | C− |
| llms.txt | Zero frontier-crawler fetches in a 7-month study | D |
| Blocking Google-Extended for AI Overview control | Does not do what people think it does | D |
Roughly 70% of AI Overview citations come from domains already ranking on page one. Pages ranking first in Google have a 43.2% ChatGPT citation rate — 3.5 times the rate of pages beyond position 20. Traditional SEO outperformed purpose-built conversational-SEO methods in the one peer-reviewed head-to-head. Relevance, recency and candidate-list position dwarfed every content-cosmetic factor in the largest controlled experiment.
That is a less exciting thesis than most AEO marketing. It is the one the data supports. And it has a practical corollary that should shape your budget: an AEO program that neglects the fundamentals of being findable is optimizing the second half of a race it has not entered.
What genuinely is new — what traditional SEO does not cover — is the entity layer, the corroboration layer, the specificity of sub-query competition, and the agent-transactability layer. Those are the five levers of Part II, and they are where the incremental return lives.
PART II — THE FIVE LEVERS
Chapter 5: Entity — Being Someone the Machine Can Name

Why this is lever one
Chapter 2 established a fact most AEO strategy skips: somewhere between two-thirds and four-fifths of prompts never trigger a web search. Google’s grounding threshold defaults to 0.3, below which the model answers from what it already knows. Profound measured web search firing on 18% of ChatGPT conversations; Semrush measured web search enabled on 34.5% of queries and falling.
For that majority of prompts, your content is irrelevant. What matters is whether the model already knows your brand exists, what category it belongs to, and what it is good at. That is not a publishing problem. It is an entity problem, and it is solved slowly, in public, across properties you do not own.
The correlational evidence points the same way. Ahrefs analyzed roughly 75,000 brands and measured Spearman correlations between AI visibility and a range of signals:
| Factor | ChatGPT | Google AI Mode | AI Overviews |
|---|---|---|---|
| YouTube mentions | 0.737 | 0.728 | 0.712 |
| Branded web mentions | 0.664 | 0.709 | 0.656 |
| Branded anchors | 0.511 | 0.628 | 0.527 |
| Branded search volume | 0.352 | 0.466 | 0.392 |
| Domain Rating | 0.266 | 0.285 | 0.326 |
| Backlinks | ~0.2–0.3 | ~0.3–0.4 | 0.218 |
| Site page count | 0.194 | — | 0.170 |
These figures combine Ahrefs’ May 2025 AI Overviews study with its December 2025 multi-platform study, so compare down a column rather than across rows.
Unlinked brand mentions correlate substantially more strongly than backlinks — roughly 0.66 against roughly 0.22. (That ratio is not a threefold difference in explanatory power; correlation coefficients do not work that way. It is a large gap, and the direction is what matters.) Content volume is nearly irrelevant. And the distribution is brutal: brands in the top quartile for web mentions averaged 169 AI Overview mentions, against 14 for the 50–75% quartile — roughly a twelvefold cliff. The bottom half of brands registered between zero and three.
Ahrefs state the necessary caveat and so will I: correlation is not causation, and there is an obvious confound. Large brands have more mentions, more YouTube presence, more branded search, and more AI visibility, all downstream of simply being large. No published study establishes causal direction. What the correlations do establish is the shape of the thing: AI visibility tracks being talked about far more than it tracks being linked to or publishing volume.
The consistency audit
Before building anything, find out what the ecosystem currently believes about you. This takes a week and it is the highest-yield diagnostic in this book, because it almost always surfaces contradictions nobody knew existed.
Assemble the canonical facts first. One page, agreed by whoever owns positioning:
Legal name, trading name, and every former name
One-sentence category statement (“X is a [category] that [does what] for [whom]”)
Founded date, headquarters location, employee band, ownership status
Primary product names and their categories
Founder and executive names, with titles
Official domain, and every domain you also own
Then audit every property against it. For each, record what it currently says and whether it matches:
| Property | Check |
|---|---|
| Homepage and About page | Category statement present in text, not just implied |
| Organization schema | Present, accurate, sameAs complete |
| LinkedIn company page | Category, founded date, HQ, employee band |
| Crunchbase | Description, founding, funding status |
| Wikidata | Item exists? Q-ID recorded? Statements referenced? |
| Wikipedia | Article exists? Accurate? (Do not edit — see below) |
| Google Business Profile | Category, NAP, hours |
| Industry directories | G2, Capterra, Clutch, trade association listings |
| Review platforms | Trustpilot, Google reviews, sector-specific |
| App stores | If applicable |
| Executive LinkedIn profiles | Do they describe the company the same way? |
| Press boilerplate | Does the boilerplate in your last ten releases match today’s positioning? |
The press boilerplate row catches more errors than any other. Boilerplate is written once and copied for years, which means a repositioning from 2024 often lives on in every syndicated release from 2019.
Score it. Count contradictions, not properties. A category described four different ways across nine properties is four contradictions. Fixing them is a project with a defined end, which makes it fundable.
Evidence grade: A−. The mechanism is documented in Google’s own generative-summary patent — confidence is derived partly from how many documents verify a claim, and low-confidence content can be suppressed. The correlational support is strong. What is unmeasured is the size of the effect from fixing a given inconsistency.
Building the Entity Home
Your homepage or About page is the canonical reference. Four requirements:
1. State identity in plain text. Somewhere in indexable HTML, in a sentence a machine can lift: “Acme Logistics is an enterprise supply chain software company serving mid-market manufacturers in North America.” Not a tagline. Not a video. Text.
Brands resist this because it reads as flat next to the brand-led copy the site was designed around. Put it in the About page’s opening paragraph if the homepage cannot carry it. But it must exist somewhere, in words, in the initial HTML response.
2. Ship Organization schema with complete sameAs. This is the one schema investment worth making regardless of Chapter 4’s findings, because its job is disambiguation, not citation. A minimal, correct implementation:
{
"@context": "https://schema.org",
"@type": "Organization",
"@id": "https://example.com/#organization",
"name": "Acme Logistics",
"legalName": "Acme Logistics Holdings, Inc.",
"url": "https://example.com/",
"logo": "https://example.com/logo.png",
"description": "Enterprise supply chain software for mid-market manufacturers.",
"foundingDate": "2016-04-12",
"address": {
"@type": "PostalAddress",
"addressLocality": "Austin",
"addressRegion": "TX",
"addressCountry": "US"
},
"sameAs": [
"https://www.wikidata.org/wiki/Q00000000",
"https://www.linkedin.com/company/acme-logistics",
"https://www.crunchbase.com/organization/acme-logistics",
"https://github.com/acmelogistics",
"https://www.youtube.com/@acmelogistics"
]
}
The sameAs array is the point. It is your machine-readable assertion that these scattered profiles are all the same entity — the thing the knowledge graph is trying to work out on its own.
3. Keep it current. A stale Entity Home teaches the graph your facts are unreliable. Put a quarterly review on someone’s calendar with a named owner.
4. Give your people entities too. If your expertise argument rests on named humans, those humans need resolvable identities: consistent bylines, author pages with credentials, author markup pointing at those pages, and consistent LinkedIn and conference-speaker profiles. An expert the system cannot identify contributes nothing to the expertise signal.
Wikidata: do this, it is the accessible one
Wikidata’s notability bar is fundamentally lower than Wikipedia’s. Wikidata admits an item that is a clearly identifiable entity describable using serious, publicly available references. Wikipedia requires significant coverage in multiple reliable secondary sources independent of the subject. Wikidata admits most real businesses. Wikipedia does not.
The process takes about an hour:
Search first, using your exact legal name. Duplicate items are common and annoying to merge.
Create an account under a real name or a clearly branded handle, and disclose any paid relationship.
Add label (business name), description (a disambiguating phrase under 250 characters), and aliases (legal name, former names, common misspellings).
Save, and record your Q-identifier. This is the durable machine-readable handle for your brand. Put it in your sameAs.
Add referenced statements: instance of, country, headquarters location, inception date, founder, official website (exactly once), industry, and external identifiers such as Crunchbase and LinkedIn IDs.
The failure mode: statements without references are the single most common reason new items get reverted. Self-published sources — your own blog, a syndicated press release — do not count. Thin entries get deleted, and a deleted item is harder to recreate than a good one is to build.
Evidence grade: A−. Low cost, documented mechanism, direct feed into knowledge graph resolution. Nobody has isolated its effect on AI citations specifically, but it is the cheapest legitimate entity work available.
Wikipedia: the warning that belongs in your agency contract
Most companies do not qualify for a Wikipedia article, and pursuing one anyway is an active reputational risk.
The rules, which every marketing leader should know before anyone on their team touches the site:
Paid advocacy is forbidden. The Wikimedia Foundation calls it a black-hat practice.
Paid editing must be disclosed — employer, client, affiliation — on a user page, on the talk page, or in the edit summary. Non-disclosure violates the Wikimedia Terms of Use.
Do not edit the article directly. The sanctioned route is to post on the talk page using the {{request edit}} template, disclose fully, and accept that the request may simply be declined.
Consequences of getting it wrong include account blocks, potential exposure under FTC guidelines and European fair-trading law, press coverage of the attempt, and permanent loss of control over the content.
Notability for companies requires significant coverage in multiple reliable secondary sources independent of the subject. Funding announcements, press releases, routine business-press coverage, and interviews with your own executives generally do not count.
An agency promising “we’ll get you a Wikipedia page” is selling you either a Terms of Use violation or a deletion discussion. Budget for Wikidata, which you can legitimately build, and let Wikipedia follow real notability if it ever arrives.
There is an irony worth noting: Wikipedia is the most-cited domain in AI answers — somewhere between 5% and 13% of ChatGPT citations depending on the study and period — while its own human pageviews fell roughly 8% year over year, which the Wikimedia Foundation attributes partly to generative AI. The most valuable source in the answer economy is being drained by it.
Google Knowledge Panel
You do not apply for one. Google generates them automatically “when there is enough information available on the open web.” You can claim one if you represent the subject, via Google’s verification process, after which you may suggest changes — which Google may decline. Google retains final say and will not act on facts it considers reasonably disputed.
What actually moves the needle, in rough order of leverage:
Consistent identity across authoritative third-party sources
A well-referenced Wikidata item
Organization and sameAs schema on your own site
Independent press coverage in genuinely reliable outlets
Verified social and business profiles
A Wikipedia article, if and only if you legitimately qualify
The entity playbook
| # | Action | Effort | Grade |
|---|---|---|---|
| 1 | Write the canonical facts page; get positioning sign-off | 1 day | A− |
| 2 | Audit all external properties against it; count contradictions | 1 week | A− |
| 3 | Fix contradictions, starting with LinkedIn, Crunchbase, press boilerplate | 2–4 weeks | A− |
| 4 | Add plain-text identity statement to homepage or About page | 1 day | A− |
| 5 | Implement Organization schema with complete sameAs | 1 day | A− |
| 6 | Create or complete a referenced Wikidata item; record the Q-ID | 1 day | A− |
| 7 | Build author entity pages for named experts; add author markup | 1 week | B |
| 8 | Update press boilerplate everywhere; brief the PR agency | 1 day | A− |
| 9 | Set quarterly consistency review with a named owner | ongoing | A− |
| 10 | Claim the Knowledge Panel if one exists | 1 day | B |
Nothing on this list is expensive. Most of it has been sitting undone in every organization I have looked at, because it belongs to nobody in particular. Assign it.
Chapter 6: Retrievability — Making Content Reachable and Readable
This is the least glamorous lever and the one with the strongest evidence behind it. It is also the one where a single misconfiguration can silently remove you from a platform entirely.
The crawler map, and the mistake almost everyone makes
The most commonly botched item in AEO strategy is confusing the bots that gate training with the bots that gate citation. Get it wrong and you either block yourself out of answers or hand over training data you meant to withhold.
Sources: OpenAI, Anthropic, Perplexity and Google crawler documentation, all first-party.
| Platform | User agent | Purpose | Effect of blocking |
|---|---|---|---|
| OpenAI | GPTBot | Training foundation models | Excludes you from training. Does not affect ChatGPT search citation |
| OpenAI | OAI-SearchBot | Powers search results in ChatGPT | Blocks you from ChatGPT search answers. This is the citation-critical one |
| OpenAI | ChatGPT-User | User-initiated fetches | Not used for automatic crawling or ranking |
| OpenAI | OAI-AdsBot | Safety-checks pages submitted as ads | No training or search impact |
| Anthropic | ClaudeBot | Training | Excludes future content from training data. Search unaffected |
| Anthropic | Claude-SearchBot | Improves Claude search quality | May reduce visibility and accuracy in Claude search results |
| Anthropic | Claude-User | User-directed fetches | Reduces visibility in user-directed responses |
| Perplexity | PerplexityBot | Surfaces and links sites in results; explicitly not used for training | Must be allowed for inclusion in Perplexity |
| Perplexity | Perplexity-User | User-initiated fetches | Perplexity’s own docs: “generally ignores robots.txt rules” |
| Googlebot | Search, Discover, News — and AI Overviews and AI Mode | Blocking removes you from Search entirely | |
| Google-Extended | Not a crawler. A token governing Gemini Apps and Vertex use | Does not affect AI Overviews or AI Mode |
Three operational notes. OpenAI publishes per-bot IP ranges and warns that robots.txt changes take about 24 hours to propagate. Anthropic publishes its bot list but cautions that IP-only blocking is unreliable because its crawlers use public cloud addresses.
And a trap worth stating plainly, because it is the most likely way a reader breaks their own site while following this chapter: robots.txt is allow-by-default, so you do not need to “explicitly allow” anything. Worse, creating a named group — User-agent: OAI-SearchBot — means that crawler reads only that group and ignores your User-agent: * rules entirely, silently unblocking whatever you had disallowed globally. The correct action is to verify these agents are not disallowed, not to add groups for them. If you do add a named group, duplicate your global disallow rules inside it.
The Google-Extended misconception is worth correcting inside your organization. There is no bot by that name fetching your pages. Googlebot crawls as it always has; the token tells Google whether the resulting content may be used to train and ground Gemini Apps and the Vertex API. AI Overviews and AI Mode are features within Google Search, so the token does not touch them. The only lever there is the Search Console opt-out toggle — and you should not pull it, for reasons in Chapter 10.
Evidence grade: A. First-party platform documentation.
The Cloudflare deadline
If you are behind Cloudflare, this section has a date on it.
On 1 July 2026 Cloudflare replaced its single AI-bot toggle with three behavioural categories: Search (collects or indexes content to answer questions about it later), Agent (acting in real time on a person’s behalf), and Training. For each you choose allow, block, or block only on pages with ads. Available to all customers including free tier.
Effective 15 September 2026, for new domains, new sites on existing accounts, and free-tier customers: on pages that display ads, Training and Agent are blocked by default; Search remains allowed.
The trap is in the interaction: multi-purpose crawlers that combine Search and Training are blocked if Training is disabled. The most restrictive applicable rule wins. A site that monetizes with ads and takes the new defaults can silently lose AI-search visibility from mixed-purpose crawlers.
Action item with a deadline: audit your Cloudflare zone security settings before 15 September 2026.
Cloudflare also introduced a content-use signal for robots.txt with three levels — immediate (no storage or reuse), reference (the default: index, excerpt, link back), and full (summarize and reproduce) — and shifted its Pay Per Crawl experiment toward compensating publishers when content appears in an AI answer rather than when it is fetched. Launch partners are small and general availability has no date. Watch it; do not plan revenue against it.
For context on why any of this exists: Cloudflare reports 57.5% of HTML web traffic was bots as of June 2026, the first time automated requests exceeded human ones. Crawl-to-referral ratios are extreme — ClaudeBot around 11,000:1 in late May 2026 (improved from roughly 24,000:1), GPTBot around 1,276:1, PerplexityBot 111:1 — against benchmarks of 4.9:1 for Google and 1.5:1 for DuckDuckGo. The economics of the open web are being renegotiated, and access control is the negotiating table.
Agentic browsers: your WAF is now a marketing decision
A new traffic class has appeared that almost nobody has written policy for. HUMAN Security’s April 2026 measurement found browser-based agents making up roughly 71% of observed agentic activity: Perplexity’s Comet at 48.1%, OpenAI’s Atlas at 21.3%, Claude’s Chrome extension at 17.3%, ChatGPT Agent at 8.6%.
Where they go: 69.6% of agentic activity touched product and search routes. Authentication was 9.2%. Checkout and payment was 3.2% — independent corroboration that agents currently browse rather than buy.
The practical risk: these agents look like browsers, not declared crawlers. Aggressive bot mitigation blocks them, and when it does you are not blocking a scraper, you are blocking a customer’s assistant mid-task. Get your security team and your marketing team in the same room about this once, deliberately, before it happens.
Render it on the server
Chapter 4 established this as the least contested finding in the field. Here is what to do about it.
Anything you want cited must be in the initial HTML response. Not after hydration. Not in a tab that loads on click. Not behind an intersection-observer lazy-load. Not in a JavaScript-injected accordion.
The first test is view-source, not the DevTools inspector. The inspector shows you the rendered DOM after JavaScript has run — which is exactly what most AI crawlers never see. The definitive test is curl with the bot’s own user-agent string, since edge logic, geo rules and user-agent-based rendering can all serve a crawler something different from what your browser receives. If your key content is not in view-source, it does not exist for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot or PerplexityBot. Google’s AI features inherit Googlebot’s rendering, and Gemini renders during live fetch. Everyone else does not.
Then fix your crawl waste. A third of AI crawler fetches on the average site hit 404s. Pull your server logs, filter by AI user agent, sort by status code, and fix the top offenders. This is a one-afternoon job that measurably increases the share of your real content that gets retrieved.
Evidence grade: A. Vercel and searchVIU independently, with no dissenting measurement.
Passage-level structure — what is actually supported
Here is where I have to be more careful than most books on this subject, because the evidence cuts against the received wisdom.
What is documented: retrieval operates on chunks of a few hundred tokens. Google’s search chunker caps at 500 tokens and offers heading inclusion explicitly to prevent context loss. Adjacent chunks can be returned alongside a match.
What is measured: the largest controlled experiment found formatting and content structure had minimal effect across all six models tested. The peer-reviewed C-SEO Bench found conversational-SEO rewrites frequently harmed ranking. One study reported that body-only optimization reduced top-20 presence by around 9%.
What follows: structure your content well because it makes retrieval mechanically plausible and because it is good writing — not because there is evidence of a large visibility lift. Specifically:
Avoid orphaning dependencies. A section that opens “As mentioned above” is ambiguous when retrieved alone. Rewrite the opening sentence to restate its subject. This is cheap and mechanically sound.
Put the answer near the top of the section. Citations peak in the 10–20% depth region of a page and fall to 2–4% in the bottom tenth. Front-load.
Match question-form headings to question-form queries. The one well-evidenced structural finding is that question-form queries trigger AI Overviews 6.8 times more often than non-question queries. That is about how users ask, not how you format — but organizing content around real questions is how you meet those queries.
Do not fund a wholesale reformatting project. Grade C. If someone proposes restructuring 400 pages into bullet lists for AI visibility, ask them for the evidence and then show them this chapter.
The Search Console opt-out: do not pull it
Google added a property-level toggle to exclude your site from AI Overviews, AI Mode and Discover AI features while remaining in classic Search. It began appearing outside UK accounts in July 2026. It does not affect rankings, Shopping, Ads, or model training.
Three reasons not to use it:
You cannot make the decision on data. Search Console reports AI impressions but not clicks, CTR, or queries. You would be trading an unmeasured benefit for an unmeasured cost.
It may take more than you intend. Reporting indicates roughly 15–17% of trending news queries embed Top Stories carousels inside AI Overviews. Opting out may remove those placements too — a consequence Google does not clearly disclose.
It is unilateral disarmament in a comparative system. Chapter 2: ranking is pairwise. Removing yourself from the candidate set does not reduce your competitor’s visibility; it increases it.
The retrievability playbook
| # | Action | Effort | Grade |
|---|---|---|---|
| 1 | Audit robots.txt against the crawler table; verify OAI-SearchBot, PerplexityBot and Claude-SearchBot are not disallowed | 1 day | A |
| 2 | Audit Cloudflare/WAF settings before 15 September 2026 | 1 day | A |
| 3 | Verify key content appears in view-source, and re-check with curl using each bot’s user agent | 2 days | A |
| 4 | Move any client-side-only primary content to SSR/SSG | varies | A |
| 5 | Pull AI-bot server logs; fix top 404s and redirect chains | 1 day | A |
| 6 | Confirm no noindex/nosnippet on pages you want in AI answers | 1 day | A |
| 7 | Review bot-mitigation rules against agentic browser user agents | 1 day | B+ |
| 8 | Rewrite opening sentences of sections that begin with back-references | ongoing | B− |
| 9 | Front-load answers within sections | ongoing | B− |
| 10 | Do not pull the Search Console AI opt-out | — | A |
| 11 | Do not fund a site-wide “chunkability” reformatting project | — | C |
| 12 | Publish llms.txt only if it costs an hour and nobody bills for it | 1 hour | D |
Chapter 7: Corroboration — Earning the Consensus
Chapter 2 established the mechanism: grounding confidence is derived partly from how many documents verify a claim, and low-confidence content can be suppressed. Chapter 5 established the correlation: unlinked brand mentions predict AI visibility substantially more strongly than backlinks. Those are two separate arguments, not one — the grounding mechanism concerns how many documents verify a claim, and the correlation concerns brand prevalence. Neither proves the other. Both point here.
This chapter is about the work that produces those mentions. It is the most expensive lever, the slowest, the least controllable, and — on the available evidence — the one with the highest ceiling.
Where AI answers actually come from
Start by fixing a number that is wrong almost everywhere it appears.
You have probably read that Wikipedia accounts for nearly 48% of ChatGPT’s citations. The underlying figure is real, from Profound’s analysis of 680 million citations. But 47.9% is Wikipedia’s share of ChatGPT’s top-ten domains only — not of all citations.
| Metric | Wikipedia | Forbes | |
|---|---|---|---|
| Share of ChatGPT’s top-10 sources | 47.9% | 11.3% | 6.8% |
| Share of all ChatGPT citations | 7.8% | 1.8% | 1.1% |
Anyone writing “48% of ChatGPT’s citations are Wikipedia” is off by a factor of six. I include this not to score a point but because it is the cleanest available illustration of how statistics in this field get laundered: a correctly reported number, stripped of its denominator, repeated until it becomes common knowledge.
Profound’s later and larger study, drawing on roughly 730,000 real ChatGPT conversations from late 2025, puts Wikipedia near 5%, Reddit near 3%, and — the number that matters more — the top ten domains accounting for only 12% of all citations, with a Gini coefficient of 0.8. Similarweb, measuring roughly 600,000 citation events in early 2026, found Wikipedia at 13.2% and Reddit at 12.0% for ChatGPT.
So across three independent measurements Wikipedia’s real share of ChatGPT citations sits somewhere between 5% and 13%. It is first or second. It is not half.
Where Reddit sits. Ahrefs, analyzing over three million US queries in mid-2026, found Reddit the single most-cited domain in Gemini at 29.2% mention share, with YouTube second at 13.9% and Wikipedia third at 12.1%. In Perplexity, YouTube led at 31.2%, Reddit second at 13.9%. Note the methodology: Ahrefs defines mention share against the summed citations of the top fifty domains, not the whole corpus. Do not put these percentages on the same chart as Profound’s or Similarweb’s.
The finding that should change your allocation comes from BrightEdge, which compared five engines across nine industries and measured overlap two ways:
Citation-source overlap between engines: 16%–59%.
Brand-recommendation overlap: 36%–55%.
Their summary: the engines disagree substantially about where to pull information from. They agree far more consistently about which brands belong in the answer.
Kevin Indig reports the same phenomenon from a different angle: 91% of citations appear on only one of ChatGPT, Perplexity, or AI Overviews. Ahrefs measures brand-level correlation between platforms at 0.749–0.821.
Read together, this is a strategic instruction: chasing per-platform citation tactics is a treadmill. Brand-level outcomes converge. Optimize the entity and the corroboration, not the platform.
Citation is not mention
Here is the finding to bring to the first meeting where someone shows you an AI visibility dashboard.
Semrush and Kevin Indig examined 3,981 domain appearances across 115 prompts, fourteen countries and four engines. 61.7% were “ghost citations” — the site linked as a source, but the brand never named in the answer. Only 13.2% got both citation and mention. Another 25.1% got a mention with no citation.
The platform split is nearly inverse:
| Platform | Cites | Mentions brand |
|---|---|---|
| ChatGPT | 87% | 20.7% |
| Gemini | 21.4% | 83.7% |
And query form matters enormously: comparative content produces 2.4 times more brand mentions than informational content, while short conversational queries generate 30 to 50 times more brand mentions than long structured prompts.
Semrush found something related earlier: only 6–27% of the most-mentioned brands in a category were also trusted as sources. In one study Zapier ranked first as a cited source and forty-fourth by brand mention.
These are two different outcomes and they need two different metrics. Being the footnote is not being the recommendation. If your dashboard reports one number called “visibility,” ask which one it is measuring.
What earns citations: the earned-media evidence
Muck Rack’s GenerativePulse is the largest study available on this question — over one million links from AI responses across ChatGPT, Claude and Gemini, July to December 2025. Muck Rack sells PR software, so it is interested research, but it is also the best-documented dataset in the category.
Earned media accounted for 82% of all citations. Non-paid media, roughly 94%. Muck Rack’s May 2026 update revised the earned-media share upward to 84%.
Journalistic sources: roughly 20–30% of all links cited.
Recency: half of all citations were published within the previous eleven months; roughly 4% within the previous seven days for Claude and ChatGPT.
Press releases grew fivefold across the six-month window; direct newswire citations rose from 0.2% to 1%.
Almost no cross-model overlap in outlets — only CNBC and Bankrate appeared frequently across more than one model.
Seer Interactive triangulates the recency finding from server logs: roughly 65% of AI bot hits target content published within the past year, 79% within two years, 89% within three, and only 6% older than six years.
Recency is a retrieval factor, and it is one of the few things on this list you can control directly. A refresh program that updates and re-dates genuinely revised evergreen content is doing retrieval work, not just housekeeping.
Third-party lists are where commercial queries are decided
For “best X” and “X vs Y” queries — the ones nearest to revenue — the evidence is unusually consistent and unusually unwelcome.
Listicles are the largest single page type by citation share at roughly 21.9%. Of those, 80.9% are third-party listicles; only 19.1% are brand-authored. Comparison pages, counterintuitively, take under 3%: engines appear to prefer pre-consolidated recommendations to side-by-side comparisons.
The concentration is severe. Roughly thirty domains capture 67% of citations within a topic; in product-comparison topics the top ten capture 46%. Indig’s phrasing is blunt: you are effectively shut out unless you build enough authority.
In B2B SaaS specifically, Foundation and AirOps analyzed 57.2 million citations across fifty companies, seven verticals and five platforms, and found: Reddit 28.0% on branded queries and 30.9% on unbranded; YouTube 19.6% and 14.9%; LinkedIn 8.5% and 15.4%; G2 at 10.8% on branded queries and out of the top fifteen entirely on unbranded. Their conclusion: brand-owned content made up only a small fraction of citations.
That G2 result is worth pausing on. Review sites appear when someone names your brand and vanish when they do not. That is a reputation-defense job, not a discovery job, and the two deserve different budgets.
The practical implication: for commercial queries, the highest-leverage work is not on your own website. It is getting accurately represented in the third-party roundups, review platforms and category comparisons that the engines actually cite. That is a digital PR and analyst-relations program, not a content program.
YouTube is the most underrated surface in this discipline
Two measurements, using different methods, point the same way:
YouTube averages roughly 20% citation share across AI platforms, and 29.5% within Google AI Overviews, where it is the single most-cited domain. AI Mode 16.6%, Perplexity 9.7%, ChatGPT 0.2%. Vimeo and TikTok sit at 0.1% each — this is not “video,” it is YouTube specifically.
Ahrefs found YouTube mentions the strongest single correlate of AI visibility across every platform tested, at r ≈ 0.71–0.74 — higher than branded web mentions, far higher than backlinks.
Note that these are different quantities — a citation share and a correlation coefficient — so they cannot literally converge. And the two datasets disagree sharply on Perplexity: BrightEdge puts YouTube at 9.7% there, while Ahrefs’ top-fifty mention share puts it first at 31.2%. That is the denominator problem from the start of this chapter, applied to a finding I happen to like. It should be flagged all the same.
Almost nobody’s AEO strategy reflects this. Causality is unestablished and the confound is obvious (big brands have more YouTube presence). But the cost of testing is low, the platform is owned by the company operating the largest answer engine, and the correlation is the strongest in the dataset.
Evidence grade: B−. Strong correlation, plausible mechanism, no causal test. Worth a funded experiment, not a bet-the-quarter reallocation.
Community participation, and the line you must not cross
UGC platforms are central to AI visibility. The temptation to manufacture presence on them is correspondingly large. Do not.
The legal position is no longer ambiguous. The FTC’s Trade Regulation Rule on the Use of Consumer Reviews and Testimonials took effect 21 October 2024, approved 5–0. It prohibits six categories:
Fake reviews and testimonials — explicitly including AI-generated ones — creating, selling, buying or disseminating them where you knew or should have known.
Reviews compensated conditional on sentiment, positive or negative.
Undisclosed insider reviews — officers, managers, employees, agents, with specific provisions covering solicitation from immediate relatives.
Company-controlled sites posing as independent review sites.
Review suppression through unfounded legal threats, intimidation, or false accusations.
Buying or selling fake social media indicators that misrepresent influence.
Civil penalties run up to $53,088 per knowing violation, and “per violation” can mean per review. That figure is a court-imposed maximum for knowing conduct, not an expected cost — but the exposure is per item, which is what makes it serious at scale.
Enforcement is live. In 2026 the FTC settled with TruHeight over employees posing as users and incentivized five-star reviews — a $4 million judgment, $750,000 payable — and filed against Premium Home Service over fake listings and fake reviews. Ten warning letters went to property managers, law firms and an accounting firm in December 2025.
Where the line actually sits, from the FTC’s own guidance — this nuance matters because the rule is often over-read into paralysis:
Paying for reviews is not per se illegal. Conditioning payment on sentiment is. “Tell us how much you loved your visit and get a $5 coupon” violates the rule. “Tell us about your visit and get a $5 coupon” does not.
Generalized solicitations to purchasers are safe, even if employees happen to respond, provided no sentiment requirement is imposed.
Employees may review, if they clearly disclose the relationship.
Sorting reviews by helpfulness is not suppression. Arranging them so negatives are unfindable may still violate the FTC Act separately.
The reputational position is worse than the legal one. Three cases:
Trap Plan / War Robots, November 2025. A marketing agency CEO publicly boasted about using fake Reddit accounts — forty-plus posts across gaming subreddits, staged as authentic player discoveries. The post was deleted within 24 hours; the agency and the publisher both issued apologies.
Rippling. Accused — and the allegations remain allegations — across multiple subreddits of astroturfing and mass-reporting critical posts. A moderator pinned a warning about “Rippling bots.” A watchdog subreddit formed. Whether or not the underlying claims are true, the content an AI now retrieves about that brand is the accusation, not the praise. That is the mechanism worth noticing: in a corroboration-weighted system, an unresolved reputational fight becomes the record.
University of Zurich, April 2025. Researchers ran undisclosed AI-generated comments on r/changemyview for months using fabricated personas. Reddit’s chief legal officer called it “deeply wrong on both a moral and legal level,” banned the accounts and issued formal legal demands to the university. The paper was withdrawn.
That last case establishes that Reddit will pursue legal remedies against undisclosed synthetic personas — against a university, which is a far more sympathetic defendant than a brand.
The asymmetry is the whole argument. The upside of manufactured UGC is a temporary bump in a metric nobody can measure reliably. The downside is a statutory maximum of $53,088 per knowing violation, a permanent negative corpus that AI systems will retrieve for years, and a press cycle. There is no version of this trade that works.
What legitimate participation looks like
In descending order of defensibility:
Answer questions in your own name with your affiliation visible. Reddit’s content policy targets repeated unsolicited promotion, not participation. A named employee giving a genuinely useful answer, with a disclosed affiliation, is within the rules and within community norms.
Host AMAs through official channels, coordinated with moderators.
Publish on YouTube. Given the correlation and citation data, this is the highest-leverage UGC surface available and it carries no astroturfing risk whatsoever.
Solicit reviews generally, without sentiment conditions, from actual customers.
Fix the substance. The Rippling case is the lesson: product and service problems become AI-visible problems, because the community discussion becomes the retrieval corpus. There is no communications strategy that outruns this.
The consensus test
The deepest point in this chapter is structural. An answer engine seeking corroboration across source types encounters three narratives about you:
Owned — your website, blog, press releases. One voice.
Earned — press coverage, analyst reports, third-party roundups. A second.
User-generated — Reddit, reviews, forums, YouTube comments. A third.
If owned and earned claim excellence while UGC reflects frustration, the model synthesizes toward the UGC signal, because its architecture favors multi-source corroboration over single-source claims. This is not a value judgment the system is making. It is the grounding mechanism doing exactly what it was designed to do.
Which means the brands that win in AI-mediated discovery are the ones where the story on the website, the story told by third parties, and the story told by actual customers all point in the same direction. That is not a marketing problem. It is an operating one, and it is the reason this lever belongs to the executive team rather than to the content calendar.
The corroboration playbook
| # | Action | Effort | Grade |
|---|---|---|---|
| 1 | Map the third-party roundups and review pages that currently rank and get cited for your top ten commercial queries | 1 week | B+ |
| 2 | Correct factual errors about you in those sources — the fastest win available | 2–4 weeks | B+ |
| 3 | Build an analyst and journalist relations program aimed at the cited outlets, not the prestigious ones | ongoing | B+ |
| 4 | Publish original research nobody else has; make the data quotable | quarterly | B |
| 5 | Fund a YouTube program with real substance | ongoing | B− |
| 6 | Institute a content refresh cycle; recency is a retrieval factor | ongoing | B |
| 7 | Establish a named, disclosed community presence on the two forums where your category is discussed | ongoing | B |
| 8 | Run a general, sentiment-neutral review solicitation program | ongoing | B |
| 9 | Brief legal and agencies on the FTC review rule; put it in every contract | 1 day | A |
| 10 | Never buy, seed, incentivize by sentiment, or generate reviews | — | A |
Chapter 8: Specificity — Winning the Sub-Query
The inversion
Traditional SEO rewarded top-of-funnel investment. You built broad category pages targeting high-volume head terms, accumulated domain authority, and let that authority lift everything beneath it. Volume first, specificity later.
Answer engines invert this, and the reason is mechanical rather than philosophical.
When a user asks a specific question, the system decomposes it into sub-queries targeting each attribute. A brand with content addressing exactly that intersection of attributes wins retrieval for that intersection, regardless of its overall domain authority. In the vector space of retrieval, you compete on proximity to a specific query embedding, not on aggregate reputation.
This is what Michael King calls relevance engineering, and it creates a genuine structural opening for specialists. A niche brand producing authoritative, data-rich content on a narrow domain achieves higher vector proximity than a generalist whose content is broad and shallow. Large brands’ legacy SEO libraries — optimized for high-volume head terms — often produce few passages that survive chunking and comparison for the specific attribute intersections that agentic queries target. They cover topics without addressing precise combinations.
But hold this against Chapter 4. Roughly 70% of AI Overview citations come from domains already ranking on page one, and pages ranking first have a citation rate 3.5 times that of pages beyond position 20. Specificity is the differentiator within the retrievable set. It is not a substitute for being in it. A specialist page that ranks nowhere and is cited by nobody is not winning a vector-space contest; it is invisible.
The honest formulation: authority gets you into the candidate pool. Specificity wins the comparison inside it. You need both, and most organizations have exactly one.
Prompts are not keywords
The single most important operational finding for content teams comes from Semrush’s clickstream analysis: 65% to 85% of ChatGPT prompts do not match any keyword in a 27-billion-keyword database.
Read that again if you run a content program built on keyword research. The majority of the demand you are trying to serve does not appear in your tooling, because it is not phrased the way search queries are phrased. Queries in AI Mode run roughly three times longer than classic search queries. People type “best CRM software”; they ask “what CRM should a twelve-person agency use if we bill on retainers and need QuickBooks sync.”
Two things follow. Your keyword tool is now a partial instrument. And the prompt space has to be constructed rather than looked up.
Building a prompt inventory
This is the core new workflow of the discipline. It takes about two weeks the first time and a day per quarter thereafter.
Step 1: Harvest real language, not invented language.
The best sources, in order of quality:
Sales call recordings and transcripts. The questions prospects actually ask, in their own words. This is the richest source in most organizations and nobody in marketing has ever read it.
Support tickets and chat logs. Post-purchase questions, which map directly to comparison and evaluation prompts.
Your site’s internal search logs. Underused, and they capture natural phrasing.
Sales objection logs. Objections are comparison prompts in disguise.
Community threads in your category — Reddit, Stack Overflow, trade forums, LinkedIn comment sections.
Search Console queries, still useful, now as one input among several.
AI-platform “people also ask” style suggestions and the follow-up questions the assistants themselves propose.
Step 2: Expand into attribute intersections.
For each core buying question, enumerate the attributes a real buyer would constrain on. For a B2B software category, that is typically: company size, industry, integration requirements, budget band, deployment model, compliance regime, team maturity, and the specific job to be done.
Then generate the intersections. Not “best project management software,” but:
“project management software for a 15-person construction firm that needs offline access”
“project management tool that works with QuickBooks and doesn’t require per-seat licensing”
“simplest project management software for a team that has never used one”
You are not writing a page for each of these. You are mapping the sub-query space your fan-out competitors are being evaluated against.
Step 3: Classify by intent and by commercial value.
Class
Example
Priority
Category-defining
“what is X”
Low — you will lose to Wikipedia
Comparison
“X vs Y”
High — commercial, and comparative content produces 2.4x more brand mentions
Constrained recommendation
“best X for [specific situation]”
Highest — the sweet spot
Implementation
“how do I do X with Y”
High — retention, expansion, and demonstrates experience
Troubleshooting
“why does X do Y”
Medium — strong for experience signals
Brand-specific
“is [your brand] good for X”
High — reputation defense; review sites dominate here
Step 4: Score for winnability.
For each high-priority prompt, check three things:
Who currently gets cited? Run the prompt across two or three platforms, repeatedly (see the sampling requirements in Chapter 10). Record the cited domains.
Do you rank on page one for the closest search-query equivalent? If not, that is your first job, per Chapter 4.
Do you have genuine information gain? Do you have data, experience, or specificity nobody else has for this question? If the honest answer is no, the page will lose the pairwise comparison no matter how it is formatted.
That third test kills more content plans than any other, and it should.
Information gain: the only content principle that survives the evidence
Chapter 4 was hard on content tactics. Formatting effects were negligible in the largest controlled experiment. Conversational rewrites were sometimes harmful in the peer-reviewed benchmark. Schema showed nothing.
One thing survived, in the studies that found anything at all: content that carries information a competing passage cannot.
In the GEO paper, the three top-performing interventions were adding quotations from recognized authorities (+41%), adding statistics (+31%), and citing sources (+27%). In the Sprinklr controlled experiment, evidence-backed claims and the presence of price information were among the factors with the largest odds ratios. Keyword stuffing went backwards.
The mechanism is the pairwise comparison from Chapter 2. Faced with two passages, a reranker prefers the one that offers something the other does not: a verifiable number, a named source, a concrete specification, a real price, a dated measurement, a documented case. A passage that restates widely available information in slightly different words offers nothing to prefer.
In practice, information gain means:
Publish real numbers. Your own data, your own benchmarks, your own pricing. Price presence was the strongest single content factor in the largest controlled testbed study of this question — bearing in mind that a synthetic two-document testbed is not a measured lift in Google.
Name your sources. Attribution is a retrieval asset, not an academic courtesy.
Date everything. Recency is a measured retrieval factor. An undated page is an old page.
Publish specifications. Missing specifications was a significant negative factor. Tables of real attributes beat prose about benefits.
Write from experience. Case studies, first-hand accounts and practitioner data carry a signal generic summary cannot fake, and as the web fills with synthetic content that signal appreciates.
Be confident. Hedged tone measured worse than confident tone. This is not licence to overclaim; it is licence to state what you know plainly.
And a caution: the C-SEO Bench adoption-rate finding means these advantages compress as competitors adopt them. Information gain is not a permanent moat. It is a moat only as long as you actually have information others do not — which is an argument for original research over content production.
The bottom-up build order
Given all of this, the sequence for a content program is close to the reverse of the traditional one.
Tier 1 — Constrained recommendation pages. The exact attribute intersections your best customers describe. Real specifications, real prices, real constraints, and honest statements about who the product is not for. These pages are unglamorous, they get no traffic in a keyword tool, and they win sub-queries.
Tier 2 — Comparison content. Including comparisons against competitors, written fairly enough that the answer engine will use them. Comparative content produces 2.4 times more brand mentions than other formats. Fairness is instrumentally optimal here: a comparison that reads as marketing gets discounted during grounding, because it contradicts the third-party corpus.
Tier 3 — Implementation and troubleshooting depth. The material that proves experience. This is also, incidentally, where community citations come from, because these are the answers people link to in forums.
Tier 4 — Original research. The most durable information-gain asset, and the one most likely to be cited by the third-party sources of Chapter 7. One well-designed annual study will outperform fifty blog posts.
Tier 5 — Category education. Last, not first. You will lose “what is X” to Wikipedia, and the traffic it earns is the traffic least likely to convert. Build it for completeness, not for visibility.
What not to do
Do not mass-produce pages for every intersection. Generated permutation content fails on information gain by construction, and Ahrefs found content volume near-uncorrelated with AI visibility (r ≈ 0.17–0.19). Fifty thin pages will lose to five substantive ones and will cost you crawl budget you are already wasting on 404s.
Do not restructure your entire library for “chunkability.” Grade C, per Chapter 4.
Do not write for the machine at the expense of the reader. The C-SEO Bench result — that conversational-SEO rewrites frequently harmed ranking — is the empirical version of a point that should already have been obvious. The systems are trained on human preference. Writing that reads as machine-directed is a detectable signal, and detectable signals get discounted.
The specificity playbook
| # | Action | Effort | Grade |
|---|---|---|---|
| 1 | Harvest real customer language from sales calls, support tickets, internal site search | 1 week | B+ |
| 2 | Build an attribute matrix and enumerate intersections | 3 days | B |
| 3 | Classify prompts by intent class and commercial value | 2 days | B |
| 4 | Score for winnability: current citations, current rank, honest information gain | 1 week | B+ |
| 5 | Publish prices and specifications wherever commercially possible | varies | B+ |
| 6 | Build Tier 1 constrained-recommendation pages for the top 20 intersections | 1 quarter | B |
| 7 | Build fair comparison content, including against competitors | 1 quarter | B |
| 8 | Commission one piece of original research per year, minimum | annual | B |
| 9 | Date and refresh everything on a defined cycle | ongoing | B |
| 10 | Do not mass-generate permutation pages | — | C− |
Chapter 9: Transactability — Being Usable by an Agent
The state of play, honestly
There is a great deal of noise about agentic commerce and a much smaller amount of production reality. Both matter, and the gap between them is where budgets get wasted.
What the forecasts say. McKinsey projects that by 2030 the US B2C retail market alone could see $900 billion to $1 trillion in “orchestrated revenue” from agentic commerce, with global projections of $3–5 trillion. Morgan Stanley projects $190–385 billion in US e-commerce spending from agentic shoppers by 2030, framed as 10% of US online retail in the likely case and 20% in the optimistic case.
Two corrections before anyone puts these on a slide. McKinsey’s figure is a range, not a point estimate, and “orchestrated revenue” is a far broader concept than agent-completed checkouts — it counts revenue where agents coordinate or influence the journey. Morgan Stanley’s $385 billion is the top of a range, and it covers US e-commerce only. Placed side by side as if they measured the same thing, they differ by roughly threefold largely because of definitional scope. They are not independent confirmations of each other.
What is actually happening. In March 2026 OpenAI abandoned native in-ChatGPT checkout and pivoted to purchases completing in retailer apps connected to ChatGPT, refocusing ChatGPT itself on product search and discovery. The stated reason: users research products in ChatGPT but do not complete purchases there, and merchant adoption of native checkout was minimal. Shopify’s president said roughly a dozen Shopify merchants were using the AI checkout tools despite integrations being available across ChatGPT, Gemini and Copilot.
Independently, HUMAN Security measured where agentic traffic actually goes: 69.6% to product and search routes, 3.2% to checkout and payment.
The synthesis: AI commerce today influences discovery, not conversion. Budget accordingly. An “agentic commerce transformation” program in 2026 is funding a future. A clean product feed is funding the present.
The protocol map
Four acronyms circulate. Here is what each is and whether it should be on your roadmap.
Protocol
Owner
Purpose
Status, August 2026
On your roadmap?
MCP (Model Context Protocol)
Anthropic-originated, open
Connects models and agents to tools and data
Spec release 28 July 2026; now stateless
Indirectly — it is the substrate
UCP (Universal Commerce Protocol)
Google, open source
Agent-to-retailer commerce, discovery through post-purchase
Live; Universal Cart shipped in the US in summer 2026
Yes, if you sell online
ACP (Agentic Commerce Protocol)
OpenAI and Stripe, Apache 2.0
Buyer–agent–merchant checkout
Beta; latest stable release 17 April 2026
Sequence behind UCP
AP2 (Agent Payments Protocol)
Agentic payments and spending mandates
Rolling out with Gemini
Watch
A2A (Agent2Agent)
Linux Foundation
Agent-to-agent interoperability
v1.0 shipped
No — enterprise interop, not marketing
UCP is the live one. Google announced it in January 2026, co-developed with Shopify, Etsy, Wayfair, Target and Walmart, and endorsed by twenty-plus partners including Adyen, American Express, Mastercard, Stripe and Visa. Google Shopping’s Universal Cart launched in the US in summer 2026 across Search and the Gemini app, with UCP-powered checkout live at Nike, Sephora, Target, Ulta Beauty, Walmart, Wayfair and select Shopify merchants. YouTube and Gmail integration and expansion to Canada and Australia are scheduled.
The single most concrete, checkable technical artifact in this entire space is UCP’s JSON manifest at /.well-known/ucp, advertising your supported services, endpoints and payment configurations so agents can discover capabilities without a hard-coded integration. Your engineering team can verify in one HTTP request whether you have one. Very few brands do.
ACP requires three pieces from a merchant: a product feed (CSV or JSON with identifiers, descriptions, pricing, inventory, media and fulfilment, refreshed regularly), agentic checkout endpoints where checkout state and payment processing stay on your systems, and delegated payment where OpenAI sets a maximum chargeable amount and expiry. Instant Checkout remains limited to approved partners, and OpenAI charges a fee on completed purchases that has never been published. Given the March 2026 retreat, treat ACP as a watching brief.
The eight attributes that actually matter right now
If you sell products, this is the highest-certainty AI visibility work available in 2026, because it is specified in first-party documentation rather than inferred.
Google connected eight Merchant Center attributes to AI Mode and AI Overviews, rolling out from January 2026:
Product Highlight — short selling benefits, 1–150 characters, 2–100 per product. Google’s language: helps customers discover products on AI-driven surfaces.
Product Detail — technical specifications by section, attribute and value.
Variant Option — differences beyond colour and size; repeatable up to thirty times.
Item Group Title — a generic title for a variant group, up to 150 characters, distinct from variant titles.
Related Product — six relationship types: part of set, required part, often bought with, substitute, different brand, accessory.
Question and Answer — up to 1,000 characters each for the question and the answer, 10,000 characters combined per product, up to thirty pairs. Google states this is “primarily intended for conversational experiences.”
Document Link — up to five PDFs such as manuals and specification sheets, which Google uses to “answer detailed questions in AI Mode.”
Popularity Rank — a 0–100 sales performance signal.
Note the irony in attribute six. FAQPage schema on your website was fully deprecated by Google in 2026 — rich results stopped appearing in May, reporting was removed in June, API support ended in August. Meanwhile Q&A attributes in your product feed are explicitly built for conversational surfaces. The channel moved, not the content need. If you have a library of good customer questions and answers sitting in deprecated FAQ blocks, they belong in the feed.
Evidence grade: A. First-party platform specification with stated AI-surface purpose.
The agent’s-eye view of your site
Beyond feeds and manifests, there is a broader question worth asking once a quarter: can an agent actually complete a task on your site?
Run the test yourself. Take an agentic browser, give it a realistic task — find a specific product variant, get a quote, check availability in a location, download a spec sheet — and watch where it fails. The failure modes are consistent and mostly mundane:
Content behind interaction. Specifications in a tab that loads on click. Pricing that requires a configurator. Availability behind a postcode form with no URL state.
Bot mitigation. The agent gets challenged and stops. Your WAF just declined a sale.
Unstable URLs. Agents re-fetch. Session-dependent URLs and query-parameter carts break on the second visit.
Information only in images or video. Specification tables rendered as JPEGs are invisible.
Ambiguous fulfilment terms. Shipping, returns and availability stated in prose that requires interpretation rather than in structured, machine-readable form.
Each of these is also a human usability problem, which is the useful part: agent-readiness work has a defensible business case even if agentic commerce underdelivers.
What to fund and what to defer
Fund now:
A clean, complete product feed. It is the common denominator across Google Merchant Center, OpenAI’s feed specification and Perplexity’s merchant program. Highest certainty investment in this chapter.
The eight Merchant Center attributes, in priority order: Product Highlight, Product Detail, Question and Answer, Document Link.
/.well-known/ucp, if you sell in a market where Universal Cart is live. This is one engineering ticket.
Structured, machine-readable fulfilment terms — price, availability, shipping, returns — in the initial HTML.
An agent-completability test once a quarter.
Defer:
ACP checkout endpoints, until OpenAI’s direction settles.
Anything sold as “agentic commerce readiness” as a programme rather than a set of tickets. The work is a feed, a manifest, and a measurement loop. If a proposal is larger than that, ask what specifically it ships.
The transactability playbook
| # | Action | Effort | Grade |
|---|---|---|---|
| 1 | Audit product feed completeness and refresh cadence | 1 week | A |
| 2 | Populate Product Highlight and Product Detail across the catalogue | 2–4 weeks | A |
| 3 | Migrate FAQ content into feed Question and Answer attributes | 2 weeks | A |
| 4 | Attach manuals and spec sheets via Document Link | 1 week | A |
| 5 | Publish /.well-known/ucp if selling where Universal Cart is live | 1 sprint | A− |
| 6 | Put price, availability and fulfilment terms in initial HTML | 1 sprint | A |
| 7 | Run a quarterly agent-completability test on core journeys | 1 day/quarter | B+ |
| 8 | Review bot mitigation against agentic browser traffic | 1 day | B+ |
| 9 | Watch ACP; do not build against it yet | — | B |
| 10 | Do not fund an “agentic transformation” programme | — | C |
PART III — RUNNING IT
Chapter 10: Measuring What Cannot Be Measured Cleanly
Start by saying the uncomfortable thing out loud
Before any dashboard, any tool selection, any board slide, someone in your organization has to say this sentence:
We are being asked to manage a channel whose cost is measurable and whose benefit is not.
Google Search Console shows AI impressions but not clicks, CTR, queries, or position. Bing Webmaster Tools shows citations but not traffic value. GA4 catches referrals but loses in-app and agent traffic to Direct. Meanwhile position-one click-through rate on AI Overview keywords is down 58%, and 68% of searches end without a click.
Any strategy that does not open by acknowledging that is selling something. Any measurement plan that promises to close the gap is overpromising. What follows is a plan to manage the gap honestly, which is the achievable goal.
The metric that replaces rank
The source paper for this book makes an argument worth preserving.
Under classic search there was one results page and one position. Under personalized, non-deterministic generation, there is no universal ranking. A brand that appears first for one user may not appear at all for another. So the relevant quantity is not rank but expected rank across the distribution of users — the average of personalized rankings across your user population.
The formalization matters less than the consequence. What actually breaks rank as a metric is not the averaging — it is the variance: across users, because of personalization, and across runs of the same prompt, because of stochasticity. A single observed position tells you almost nothing about either.
That is the right conceptual target. In practice you cannot observe the user distribution, so you approximate it with three things: a prompt set that represents real demand, enough repeated runs to average out stochasticity, and enough platform coverage to represent where your buyers actually are.
Miss any of the three and you are producing noise with error bars you cannot see.
The sampling requirement nobody meets
The University of St. Gallen study is the most consequential measurement research in this field and it is barely cited by the industry it indicts.
Running identical prompts on identical days produced source overlap of Jaccard 0.34–0.42 and brand-mention overlap of 0.45–0.59. The instability persisted within 24-hour windows, so it is model stochasticity, not news churn.
Their recommendations:
At least seven runs per prompt per day for brand tracking, to bring standard error below 0.10
At least eight runs for source coverage
Rolling windows of two to four weeks
Their conclusion: single observations are misleading.
Most commercial tools sample far below this and almost none publish confidence intervals. So three rules follow, and they should be non-negotiable in your organization:
Never report a single-day share-of-voice figure. Report rolling averages over two to four weeks only.
Ask every vendor, in procurement, how many runs per prompt per day they execute and whether they publish confidence intervals. The answer is diagnostic. Most will not have one.
Treat any before-and-after case study without disclosed repeated sampling as unproven. Including your own. Especially your own.
Two families of tool, and why they disagree
Every vendor sells “share of voice in AI answers.” They do not measure the same thing.
Family A — synthetic prompting. The tool maintains a prompt list you configure, fires it at model APIs or scraped chat interfaces on a schedule, parses responses for brand mentions and cited URLs, and aggregates. Peec, Otterly, Scrunch, Rankscale, Athena, Conductor, BrightEdge and Semrush all work substantially this way. Strength: you control the prompt set and get per-prompt diagnostics. Weakness: you are measuring answers to questions you invented.
Family B — real-prompt corpora. Ahrefs Brand Radar and Profound’s Index claim to derive prompts from observed behaviour rather than invention. Ahrefs states it indexes over 473 million monthly prompts across six platforms; Profound claims an index built from over 1.5 billion real user conversations, clustered and reduced to representative prompts.
The honest caveat on Family B: neither publishes an auditable methodology for how they obtain real prompts at that scale. Treat volume claims as directional, not audited.
A working selection: buy one Family B tool and one Family A tool. They will disagree. The disagreement is the methodology, not a bug, and seeing both keeps you honest.
Approximate public pricing as of August 2026, which is useful mainly for scale:
| Tool | Family | Entry price | Note |
|---|---|---|---|
| Profound | Hybrid | $99/mo starter, $399 growth | Starter is ChatGPT only; adds server-side agent analytics |
| Ahrefs Brand Radar | B | $199/mo, $699 all platforms | Real-prompt index; bundled with Ahrefs |
| Peec AI | A | Tiered; third-party listings suggest €89–499 | Prices not rendered on vendor pricing page — verify |
| Otterly | A | $29 lite, $189 standard, $489 premium | Claude, AI Mode and Gemini are paid add-ons |
| Scrunch | A | $250–500/mo | Pairs answer monitoring with crawler observability |
| Semrush AI Visibility | A | $99/mo annual, per domain | Base tier is 25 custom prompts — a tripwire, not a system |
| Rankscale | A | $20–780 credit-based | 17+ engines |
| Conductor / BrightEdge | A | Enterprise, unpublished | Conductor exposes fan-out sub-queries; BrightEdge covers only three engines |
Prices change. Verify before quoting. The structural point is that a serious measurement programme costs a few thousand dollars a year, not a few hundred thousand — which means the constraint is analytical discipline, not budget.
The free first-party sources you are probably ignoring
Google Search Console generative AI reports, launched 3 June 2026. You get impressions from AI Overviews and AI Mode, segmented by page, country, device and date. An impression counts when a link to your site is shown inside a generative AI feature.
You do not get clicks, CTR, queries, average position, or conversions. Google says it is still working out which additional metrics would be useful. A regulatory clock is running: the UK CMA’s interpretive notes require impressions, click-throughs and CTR, with additional obligations reported from December 2026 and page-level AI controls due March 2027.
One trap: impression counts differ by aggregation level. At property level, several URLs from your site inside a single AI response may count as one impression; at page level they may count separately. Do not compare property-level and page-level totals.
Bing Webmaster Tools AI Performance report, public preview since 10 February 2026. Total citations, average unique URLs cited daily, sample triggering queries, per-page citation counts, trends across Copilot and Bing AI summaries. Free, and the only first-party citation count anyone publishes. Same limitation: no clicks, no traffic value.
GA4’s AI Assistant channel group, added May 2026. When GA4 detects a referrer from a recognised AI assistant it sets medium to ai-assistant. Three limitations to plan around: traffic from in-app browsers and mobile apps often arrives with no referrer and lands in Direct, so your AI number is a floor; Google has not published the recognised-referrer list; and it measures referrals, not influence — the dominant effect is answers where nobody clicks.
The magnitude of that first limitation is measurable. One nine-month first-party study tracking 51,200 AI Overview click events found 22.4% of AI Overview traffic misattributed to Direct rather than Organic Search. Whatever your analytics says about AI traffic, the real figure is higher.
Server and CDN logs. The most underused source in the stack. Filter by AI user agent and you get crawl frequency by bot, which content is being fetched, 404 and redirect waste, and whether your robots.txt and WAF changes actually took effect. This costs nothing and answers questions no dashboard will.
Designing a prompt set that is worth measuring
In Family A, prompt-set design is the entire ballgame. A tool tracking 25 prompts is a tripwire, not a measurement system.
A workable structure for a mid-sized programme, roughly 120–200 prompts:
Rules:
Source them from Chapter 8’s harvest, not from a brainstorm. Real customer phrasing, not marketing phrasing.
| Segment | Share | Purpose |
|---|---|---|
| Constrained recommendation (“best X for [situation]”) | 40% | Commercial core |
| Comparison (“X vs Y”, “alternatives to Y”) | 20% | Competitive position |
| Brand-specific (“is [brand] good for X”) | 15% | Reputation defense |
| Implementation and troubleshooting | 15% | Experience signals |
| Category education | 10% | Coverage baseline |
Include prompts you currently lose. A prompt set built only from your strengths produces a dashboard that flatters you and teaches you nothing.
Include your top three competitors’ strongest positions.
Freeze the set for a quarter. Changing prompts mid-quarter destroys comparability, and the temptation to swap out losers is strong.
Version it. When you do change it, record the change and treat the series as broken at that point.
What to report, and to whom
Four metrics, each answering a distinct question. Do not collapse them into a single “AI visibility score” — those scores hide the trade-offs that matter.
1. Mention rate. Share of tracked prompts, on a two-to-four-week rolling basis, in which the brand is named in the answer. This is the recommendation metric.
2. Citation rate. Share of tracked prompts in which your domain is cited as a source. Remember the ghost-citation finding: 61.7% of appearances are citations without mentions, and ChatGPT and Gemini are nearly inverse on this. Reporting one and calling it the other is the most common error in this field.
3. Sentiment and accuracy. What is being said, and whether it is correct. This is the metric that catches problems before they become Chapter 12 problems, and it needs a human reading a sample of answers, not a classifier score.
4. Competitive share. Your mention rate against your named competitors’ on the same prompt set. In a comparative system, the absolute number matters less than the relative one.
Alongside those, report the inputs you control — entity consistency score, share of key content server-rendered, AI crawler 404 rate, third-party roundups where you appear correctly, feed attribute completeness. Inputs move before outputs do, and they are the only part of the system where cause and effect are legible.
The ROI conversation
You will be asked what this returns. Here is how to answer without lying.
What you can attribute: AI referral sessions and their conversion behaviour, from GA4 plus server logs, understanding the figure is a floor. If your AI referral traffic converts materially better than other channels, as Adobe’s +42% and multiple first-party reports suggest, that is a real, defensible, if small number.
What you cannot attribute: the influence of appearing in answers nobody clicks. This is most of the value and there is no honest way to measure it directly today.
What you can do about the gap: three things.
Proxy it with mention rate against competitors. If you are named in 40% of relevant answers and your closest competitor is named in 15%, that is a share-of-consideration position with real economic content, even if you cannot price it.
Instrument the qualitative side. Add “Where did you first hear about us?” and “Did you research this with an AI assistant?” to your demo request and post-purchase surveys. Self-reported attribution is imperfect and it is currently the only line of sight into zero-click influence. Report it in percentages, tracked over time.
Be explicit about the asymmetry in the business case. The cost of the entity and retrievability work in Chapters 5 and 6 is small and largely one-time. The cost of the corroboration work in Chapter 7 is real, but it is largely spend you were already making on PR and content, redirected. You are not asking for a new budget line proportional to an unmeasurable return. You are asking to reallocate an existing one toward mechanisms that are documented.
That last framing is what gets this funded. A request for new money against an unmeasurable return fails. A request to redirect existing money toward better-evidenced mechanisms succeeds.
The measurement playbook
| # | Action | Effort | Grade |
|---|---|---|---|
| 1 | Enable GSC generative AI reports and Bing AI Performance report | 1 hour | A |
| 2 | Verify GA4 AI Assistant channel and document its known gaps | 1 day | A |
| 3 | Set up AI bot log analysis from server or CDN logs | 1 week | A |
| 4 | Build a 120–200 prompt set from real customer language | 1 week | B+ |
| 5 | Buy one Family B and one Family A tool; ask both for run counts | 2 weeks | B |
| 6 | Report only two-to-four-week rolling averages | ongoing | A |
| 7 | Separate mention rate from citation rate in all reporting | ongoing | A |
| 8 | Add AI-assistant questions to demo and post-purchase surveys | 1 day | B |
| 9 | Track controllable inputs alongside outputs | ongoing | B+ |
| 10 | Never report a single-run result or an unversioned prompt set | — | A |
Chapter 11: The Budget Argument and Who Owns It
What the 60/40 rule actually says
The argument you will hear is that AI search requires shifting budget from performance to brand. The evidence for that is better than the evidence for most claims in this field, but the canonical citation is weaker than its users admit and you should know both halves.
Origin. The 60:40 brand-to-activation ratio comes from Les Binet and Peter Field’s The Long and the Short of It, published by the IPA in November 2013, drawing on the IPA Effectiveness Databank. Its authors framed it as a guideline for the average case, not a law.
The critique, which is legitimate and unrebutted on the merits. Byron Sharp of the Ehrenberg-Bass Institute has called the rule “very misleading,” built on “a very weird data set” — the IPA Databank consists of award submissions, meaning self-selected, self-reported campaigns that entrants believed had worked. His argument is that the number exists largely because people wanted a number.
James Hurman’s response is the honest middle: the rule may be less scientific than Sharp would like, but few would argue with the observation that marketers tend to underinvest in brand and overinvest in performance.
It varies by context, and its own authors say so. Effectiveness in Context (2018) exists precisely to demonstrate this. Optimal brand share ranges roughly from 20% to 80% depending on the nature of the purchase decision, the mechanics of purchase, product innovation, brand size and development strategy. For B2B specifically, the LinkedIn B2B Institute’s analysis of IPA Databank B2B cases recommends 50/50, while flagging its own limitations — the authors describe their findings as “tentative” with small sample sizes skewed toward UK and large-budget cases.
So the defensible claim is directional, not arithmetic. Most organizations underinvest in brand. The right ratio for your category is not 60/40 because a book said so; it is somewhere in a wide range that depends on how your customers buy.
What is actually happening to budgets
The movement in 2025–26 is away from brand investment, not toward it — which makes the argument in this chapter contrarian rather than fashionable.
NIQ’s CMO Outlook for 2026, surveying over 250 senior marketing decision-makers, found only 55% allocating 60% or more to long-term brand building, down from 59% the prior year, and only 69% saying their C-suite believes in long-term brand value, down from 80%. Meanwhile 74% report heightened ROI scrutiny and 84% cite marketing ROI as the primary allocation metric.
Gartner’s 2026 CMO Spend Survey, covering 401 CMOs, found marketing budgets essentially flat at 7.8% of revenue, with 15.3% of marketing budgets allocated to AI, 70% saying AI leadership is critical for 2026, and only 30% reporting mature AI readiness. Gartner reports no discrete GEO or AEO line item.
So the environment is one of tightening accountability pressure at precisely the moment when the discovery mechanism is becoming less measurable. That tension is the central budget problem of this chapter, and pretending otherwise will not help you win the argument.
The actual argument for reallocation
Do not lead with the 60/40 rule. Lead with the mechanism, which is stronger.
Argument one: a growing share of prompts never trigger retrieval. Somewhere between two-thirds and four-fifths of AI queries are answered from the model’s existing knowledge. For those, brand prevalence in the training corpus and in the model’s parametric memory is the visibility mechanism. There is no content tactic that reaches them. This is a brand argument grounded in architecture rather than in effectiveness-award data.
Argument two: you cannot buy your way in. Using classic search as the reference, paid placement and organic recommendation are structurally separated, because introducing paid bias into organic recommendations erodes the platform value that makes the two-sided market work — especially in a competitive environment where users have shown high willingness to switch platforms. ChatGPT’s share of generative AI web traffic fell from roughly 76% to 53% in twelve months. Platforms in that position cannot afford to sell their recommendations.
The counterweight, which you should state rather than hide: ads are already appearing inside answers. OpenAI began testing them in February 2026 and by mid-2026 roughly 26% of ChatGPT responses contained one, labelled and visually separated. The organic and paid layers are separating, exactly as they did in search. What remains unlikely is paid influence over the organic recommendation.
Argument three: the correlational evidence points at brand signals, not content signals. Unlinked brand mentions correlate substantially more strongly with AI visibility than backlinks — about 0.66 against 0.22 — and content volume is nearly uncorrelated. Correlation is not causation and the confound is real. But if you are allocating under uncertainty, allocate toward the signal with the stronger relationship.
Argument four — the one that actually closes. You are not asking for a new budget line. You are asking to redirect existing spend toward better-evidenced mechanisms:
Some content production budget moves to original research, which serves both information gain and earned media.
Some link-building budget moves to digital PR aimed at the outlets that actually get cited, which produces mentions rather than links.
Some technical SEO budget moves to server-side rendering and crawl-waste cleanup, which is cheaper than most of what it replaces.
Entity and feed work is small, largely one-time, and mostly unowned today.
Where the money should go
A defensible allocation for a mid-sized programme, with the caveat that Search Engine Land’s published version of this split is an editorial recommendation rather than a measured benchmark:
| Area | Share | Rationale |
|---|---|---|
| Core search fundamentals | ~40% | Chapter 4: getting retrieved dominates everything after |
| Digital PR and earned media | ~25% | Strongest correlate in the brand-visibility data; separately, corroboration count is documented as a grounding input |
| Measurement and reporting | ~20% | The sampling requirements are real and unmet |
| Enablement and training | ~10% | Most of your team’s mental model is three years out of date |
| Experimentation | ~5% | YouTube, agent-readiness, whatever the next surface is |
The line that will get challenged is the 20% on measurement. Defend it with Chapter 10: without adequate sampling you cannot tell whether any of the other 80% worked, and you will spend the difference arguing about noise.
Who owns this
The most common organizational failure in this field is that AEO gets assigned to whoever owns SEO, and roughly half the work is not SEO work.
Map it against the five levers:
| Lever | Natural owner | Common failure |
|---|---|---|
| Entity | Brand and corporate communications | Assigned to SEO, who cannot change LinkedIn or press boilerplate |
| Retrievability | Engineering, with SEO | Assigned to SEO, who cannot change the rendering architecture |
| Corroboration | PR, analyst relations, community | Assigned to content, who cannot place earned media |
| Specificity | Content, informed by sales | Built from keyword tools instead of customer language |
| Transactability | Ecommerce and product | Nobody owns it; the feed belongs to whoever built it in 2019 |
The practical structure that works in most organizations is not a new team. It is a named owner with a quarterly forum: one accountable person, a standing session with engineering, PR, content, ecommerce and legal, and a single dashboard that all of them see. The forum exists because the failure mode is not incompetence, it is that the five levers sit in five reporting lines and nobody has the whole picture.
Legal belongs in the room. Chapter 7’s FTC rule, Chapter 12’s liability developments, and the EU AI Act’s transparency obligations are all live, and all of them are cheaper to handle as constraints on the plan than as responses to an incident.
Briefing an agency
If you are buying help, the following questions separate the substantive from the performative:
“What is your evidence that this works?” Ask for it in writing. If the answer is the GEO paper, ask whether they have read C-SEO Bench.
“How many runs per prompt per day does your measurement use, and do you publish confidence intervals?”
“Are you proposing anything that requires editing Wikipedia?” If yes, end the conversation.
“Are you proposing anything involving reviews, forum posts, or testimonials we would not want printed in a trade publication?” If yes, end the conversation.
“What proportion of the work is on our own site?” If it is most of it, they are selling you Chapter 6 and calling it Chapter 7.
“What do you think llms.txt does?” A cheap and reliable diagnostic.
“What would make you tell us this isn’t working?” Anyone who cannot answer this has no falsifiable model.
The budget playbook
| # | Action | Effort | Grade |
|---|---|---|---|
| 1 | Lead the internal case with the retrieval-threshold mechanism, not the 60/40 rule | — | B+ |
| 2 | Frame the ask as reallocation, not new budget | — | — |
| 3 | Name a single accountable owner | 1 day | — |
| 4 | Stand up a quarterly cross-functional forum including legal | ongoing | — |
| 5 | Move some content budget to original research | ongoing | B |
| 6 | Move some link-building budget to earned mentions | ongoing | B+ |
| 7 | Protect the measurement line at roughly 20% | — | A |
| 8 | Run agencies through the seven questions above | 1 day | — |
Chapter 12: The Risk Register
Every chapter so far has been about getting into the answer. This one is about what happens when the answer is wrong, when someone manipulates it, or when a regulator arrives. These are not hypotheticals; each item below has a documented case attached.
Risk 1: The system states false things about you
This is the most likely risk to materialize and the least prepared-for.
The Washington University audit of 7,583 AI Overviews verified 98,020 atomic claims against their cited sources and found 11.0% inconsistent — 4.1% actively contradicted by the source, 7.0% simply not present in it. That is a baseline error rate on claims that carry a citation, which means they look verified.
Applied to your brand, this manifests as invented product specifications, wrong pricing, wrong availability, misattributed features, discontinued products presented as current, and confident statements about your policies that do not exist.
A real case: in April 2025, Cursor’s own AI support bot invented a company policy — that users could not run sessions on multiple machines. No such policy existed. Users cancelled subscriptions over a rule that was never written. The co-founder had to state publicly: “We have no such policy.”
What to do:
Monitor for factual accuracy, not just presence. Chapter 10’s third metric. Somebody reads a sample of answers about you every month.
Make the correct facts easy to find and hard to contradict. This is the entity work of Chapter 5, and accuracy is the second reason to do it.
Establish a correction path. Know in advance who reports an error, to which platform, through which channel, and who signs off.
Fix the upstream source. Most hallucinated brand facts are not invented from nothing; they are drawn from a stale third-party listing, an outdated press release, or a competitor’s comparison page. Find it and correct it.
Risk 2: Liability for what the machine says — moving fast, in different directions
The most significant development, and it favors brands — with a caveat. On 28 May 2026 the Regional Court of Munich granted a temporary injunction against Google after AI Overviews falsely connected two publishers to scams and fraudulent practices, fabricating claims not even made in the underlying search results. The court held Google directly liable, rejecting the argument that users should verify independently, and drew an explicit line: an AI overview “is its own content, not just a list of search results.” Since only Google controls the algorithms, Google owns accuracy.
The caveat matters. This is an interim injunction from a court of first instance, and Google has said it will appeal. It is not settled precedent, in Germany or anywhere else. What it establishes is that the argument works in at least one major jurisdiction — which is a meaningful shift from where this stood a year ago, and less than the headlines claimed.
The US position is currently the opposite. In Walters v. OpenAI, a Georgia court granted summary judgment to OpenAI in May 2025 over ChatGPT fabricating embezzlement allegations, on three independent grounds: a reasonable reader in context could not have understood the output as communicating actual facts; no negligence or actual malice; no recoverable reputational harm.
But it is not settled. In Starbuck v. Google, filed October 2025 over allegedly fabricated criminal accusations and invented court documents, Google’s motion to dismiss was denied and the case advanced to discovery.
And you are liable for your own chatbot. In Moffatt v. Air Canada (BC Civil Resolution Tribunal, 14 February 2024), Air Canada’s chatbot told a passenger he could apply retroactively for bereavement fares, contradicting the website. Air Canada argued the chatbot was a separate legal entity responsible for its own statements. The tribunal rejected this outright and held the airline accountable for all information on its website, static or generated.
What to do: if you deploy a customer-facing AI assistant, treat its outputs as your published statements, because that is what a tribunal has already called them. Constrain it to verified sources, log what it says, and label it.
Risk 3: Manipulation, and why you should not participate
Adversarial manipulation of AI recommendations is demonstrated, effective, and a trap.
ETH Zürich researchers introduced preference manipulation attacks — crafted web content or plugin documentation that steers LLM-powered search toward the attacker’s product. Manipulated fictional cameras became 2.5 times more likely to be recommended, competing successfully against real brands. Fake products moved from a 34% to a 59.4% recommendation rate. Malicious plugins achieved up to 7.2 times higher selection rates. Production systems affected included Bing Copilot, Perplexity, GPT-4 and Claude 3.
Now the finding that should end the discussion internally: when multiple competitors attack simultaneously, everyone’s recommendation rate degrades. It is a prisoner’s dilemma. Individually rational, collectively destructive.
That is the argument to use with anyone in your organization tempted by this, because it does not depend on them sharing your ethics. The tactic destroys the value of the surface it exploits, and it does so quickly once more than one player adopts it — the same adoption-rate dynamic C-SEO Bench measured for legitimate tactics.
The related security exposure is real too. Brave’s researchers demonstrated indirect prompt injection in Perplexity’s Comet browser, disclosed August 2025. The proof-of-concept vector was a Reddit comment with instructions hidden in a spoiler tag: when a user asked the browser to summarize the page, the AI executed the hidden instructions, accessed the user’s account, and exfiltrated an email address and a one-time password.
Your brand’s community pages, review sections and user-generated content are now an attack surface that can be used against your customers’ assistants. Nobody’s content moderation policy accounts for this yet. Yours should.
Risk 4: Regulatory
EU AI Act, Article 50 — in force since 2 August 2026. The transparency obligations apply now, not in some future phase:
Users must be clearly informed they are dealing with an AI system unless obvious from context.
Synthetic content, including text, must be marked in machine-readable format. Systems already deployed have until 2 December 2026.
Deployers must inform individuals when emotion recognition or biometric categorisation is operating.
Penalties reach €15 million or 3% of worldwide turnover.
The direct marketing relevance: AI-generated marketing text distributed in the EU falls within the machine-readable marking obligation, and AI chat interfaces on brand properties require disclosure. High-risk obligations were postponed by the AI Omnibus to 2027 and 2028; the transparency obligations were not.
FTC “Operation AI Comply.” Launched September 2024 and continuing across a change of administration, with more than a dozen cases associated with AI washing in the past year. Two features matter for marketing leaders: enforcement has expanded to B2B marketing — the same substantiation standards apply regardless of audience — and the FTC has invoked the “means and instrumentalities” doctrine to hold vendors liable for supplying deceptive marketing materials used downstream. Your agency and your martech supplier are in scope, not only you.
FTC review rule. Covered in Chapter 7. Up to $53,088 per knowing violation, live enforcement since 2026.
Risk 5: Platform dependency
The structural risk that underlies all of the above.
Answer engines function as intermediaries between you and your customers, and their ranking logic is opaque in a way search never quite was. Under classic SEO, ranking factors were partially understood and outcomes were auditable — you could see the results page. In an agentic regime you get impressions without clicks, citations without traffic, and recommendations you cannot observe.
Three specific exposures:
Concentration. Roughly thirty domains capture 67% of citations within a topic; the top ten capture 46% in product-comparison topics. Large publishers took 82% of publisher names volunteered by AI assistants and 97% of follow-through visits. The breadth picture is more mixed than it is usually reported: the number of unique domains receiving ChatGPT referrals rose from roughly 71,000 in October 2024 to a peak near 260,000 in October 2025, then fell back to about 170,000 by February 2026 — well off its peak, but still more than double a year earlier. Treat the concentration risk as real at the topic level, and the referral-breadth trend as unresolved.
Volatility. ChatGPT’s share of generative AI web traffic fell from roughly 76% to 53% in twelve months while Gemini tripled. Any strategy tuned to one platform’s behaviour is tuned to a moving target.
Commercial encroachment. Ads inside answers, at roughly 26% of ChatGPT responses by mid-2026. The organic surface is smaller than the page it replaced and it is being monetized.
What to do: diversify deliberately. The entity and corroboration levers are the platform-independent ones — they work across engines because, per BrightEdge, engines disagree about sources and agree about brands. Owned audience relationships — email, community, direct — are the only fully independent asset. Nothing here argues for abandoning AI visibility; it argues against building a business on top of a single opaque intermediary, which is a lesson this industry has already learned once.
The risk playbook
| # | Action | Frequency | Grade |
|---|---|---|---|
| 1 | Monthly accuracy audit: read a sample of AI answers about your brand | monthly | A |
| 2 | Maintain a documented correction path per platform | ongoing | B+ |
| 3 | Trace and fix upstream sources of recurring factual errors | ongoing | B+ |
| 4 | Treat your own AI assistant’s output as published statements; log and constrain it | ongoing | A |
| 5 | Brief legal on FTC review rule, Operation AI Comply, EU AI Act Article 50 | once, then annually | A |
| 6 | Mark AI-generated marketing content machine-readably for EU distribution | now | A |
| 7 | Prohibit manipulation tactics in policy and in agency contracts | once | A |
| 8 | Add prompt-injection review to UGC moderation policy | once | B |
| 9 | Build owned-audience assets as platform-independent insurance | ongoing | A |
Chapter 13: The Next Two Years
Forecasting in this field has a poor record, so what follows is structured as scenarios with trigger conditions rather than predictions. The useful question is not “what will happen” but “what would I do differently if this happened, and what would tell me it was happening?”
Scenario 1: Hyperspecialization and segmentation
The thesis. Because retrieval rewards proximity to specific query embeddings rather than aggregate reputation, small specialists gain a structural advantage over generalist incumbents. Small brands double down on niches. Large brands respond by segmenting into sub-brands, each specialized enough to win its own attribute intersections. Either way, a new layer of specialized brands emerges between consolidated manufacturers and consumers.
The consequence, if it holds, is significant: brand equity becomes localized to specific markets rather than functioning as a portable asset that lets a large company enter new categories cheaply. That would invert one of the oldest assumptions in marketing strategy.
The counter-evidence, which is substantial. The largest observational study of brand visibility across AI platforms found a steep incumbency gradient: household names appeared in roughly 73% of relevant answers, mid-market brands 44%, small and niche brands 11%. Ahrefs’ correlation data shows the same cliff — the bottom half of brands register between zero and three mentions. Whatever theoretical advantage specialists hold in vector space, incumbents are winning in practice today.
What would tell you the thesis is winning: specialist brands appearing in your category’s constrained-recommendation prompts while absent from broad category prompts; the incumbency gradient flattening in successive studies; sub-brand launches by large competitors framed around narrow use cases.
What to do either way: the specificity work in Chapter 8 is the hedge. It costs the same whether you are the incumbent or the challenger, and it is the mechanism by which either wins a sub-query.
Scenario 2: Paid influence over organic recommendations
The question. Will companies eventually be able to pay for preferential treatment inside AI recommendations, as they bid for search advertising?
The argument against. Using search as the reference, paid and organic remain structurally separate because introducing paid bias into organic results erodes the platform value that makes a two-sided market function. Answer engines retain that structure, and they operate in a more competitive environment with demonstrated user willingness to switch platforms. Selling the recommendation would be selling the reason people use the product.
The argument for caution. The separation is already thinner than it looks. Ads appear inside roughly a quarter of ChatGPT responses. Google is piloting Direct Offers, a Google Ads programme letting retailers surface exclusive discounts inside AI Mode, announced January 2026 with Petco, e.l.f. Cosmetics, Samsonite and Rugs USA. The distinction that matters is not “ads versus no ads” — that battle is over — but whether the organic recommendation can be influenced by spend.
What would tell you it is happening: unlabelled placement changes correlated with spend; platforms offering “recommendation visibility” products; a measurable correlation between advertising spend and organic mention rate that survives controlling for brand size. That last one is a research project you could actually run.
What to do: measure organic mention rate separately from any paid placement, from the beginning. If the surfaces blur later, you will want the clean baseline.
Scenario 3: The transaction layer arrives, or does not
The bull case is McKinsey’s $900 billion to $1 trillion in orchestrated US B2C retail revenue by 2030 and Morgan Stanley’s $190–385 billion in agentic US e-commerce spending, with UCP live at major retailers and AP2 rolling out.
The bear case is what actually happened in March 2026: OpenAI retreated from native checkout after roughly a dozen Shopify merchants adopted it, because users research in the assistant and buy elsewhere. HUMAN Security’s measurement backs this — 3.2% of agentic activity touches checkout.
The likely middle, and the configuration the source paper argues is most viable, is AI initiation rather than AI transaction: the agent assembles the cart or populates checkout, and a human confirms. This is functionally equivalent to one-click purchase on your own site, preserves user authority, and avoids the accidental-purchase problem — where users may remain liable under existing electronic fund transfer regulation for transactions they never explicitly confirmed. Full autonomy creates logistical and regulatory problems that nobody has solved.
What would tell you the layer is arriving: checkout share of agentic traffic rising above single digits; UCP checkout expanding beyond the initial retailer set; a major platform shipping autonomous purchase with a published liability framework.
What to do: the Chapter 9 fund-now list, and nothing beyond it. A feed, the eight attributes, a manifest, and a quarterly agent test. That work pays off in discovery today regardless of whether transaction volume ever materializes.
Scenario 4: Measurement standardizes, or fragments further
The pressure is regulatory. The UK CMA’s interpretive notes require impressions, click-throughs and CTR; Google currently supplies only impressions, with additional obligations reported from December 2026 and page-level AI controls due March 2027. If that holds, Google will be compelled to disclose more than it has chosen to.
The counter-pressure is architectural. Personalization makes a universal ranking incoherent. Non-determinism makes single measurement meaningless. No regulator can mandate a number that does not exist.
The likely outcome is partial: better first-party impression and citation reporting, no standardized share-of-voice metric, and continued vendor disagreement. Which means the discipline of Chapter 10 — repeated sampling, rolling windows, separating mention from citation — stays a competitive advantage rather than becoming table stakes.
What would tell you standardization is arriving: Google shipping clicks and CTR in Search Console; an industry body publishing a sampling standard; vendors beginning to publish confidence intervals.
Scenario 5: The adversarial equilibrium
Model developers want accurate reviews and verifiable sources. Marketers are incentivized to bias reviews favorably and inflate credibility. This competition drives much of the innovation in the field, and it resolves in one of two ways.
Path one: brands align public sentiment with their desired image. If what customers say matches what the brand claims, there is nothing to manipulate. Research on corporate reputation finds public sentiment is “sticky” — cognitive biases resist change — but firms willing to invest in genuine change can move it. This is the expensive path and the durable one, and it is the operating-not-marketing point from the end of Chapter 7.
Path two: platforms adopt manipulation-resistant validation — independent third-party verification that brands cannot influence directly. This would end the arms race by removing the lever. No such system currently exists, and building one is a substantial trust and governance problem.
What to do: assume path one. It is the only one you control, and it is the only strategy that improves your business as a side effect of improving your visibility.
What to watch, concretely
Six indicators worth a quarterly glance:
Retrieval rate. The share of prompts triggering web search — currently 18% to 35% depending on the measurement, and falling in ChatGPT. If it keeps falling, the entity lever grows and the content lever shrinks.
Citation concentration. Unique domains receiving referrals — off its October 2025 peak but above a year earlier. If it resumes narrowing, the specialist thesis is losing.
Ad density inside answers. Currently around 26% of ChatGPT responses.
Checkout share of agentic traffic. Currently 3.2%.
Platform share volatility. ChatGPT down roughly 23 points in a year; Gemini up roughly 19.
First-party reporting. Whether Google ships clicks and CTR.
Each of these has a threshold at which your allocation should change. Write those thresholds down now, while you are not under pressure, and revisit them quarterly.
Chapter 14: The First 90 Days
Everything in this book, sequenced. Dependencies run downward: entity before corroboration, retrievability before content, measurement before claims of success.
Effort estimates assume a mid-sized organization with an existing SEO function, and are sequenced here rather than sized as in the chapter playbooks — a task shown as “one quarter” in a playbook may appear as a six-week slot here because it starts inside the ninety days rather than finishing inside them. Evidence grades carry over unchanged.
Days 1–30: Diagnose and stop the bleeding
The goal of the first month is to know what is true and to fix the things that are silently broken.
| # | Action | Owner | Effort | Grade |
|---|---|---|---|---|
| 1 | Write the canonical facts page; get positioning sign-off | Brand | 1 day | A− |
| 2 | Audit every external property against it; count contradictions | Brand/SEO | 1 week | A− |
| 3 | Audit robots.txt against the crawler table; verify OAI-SearchBot, PerplexityBot and Claude-SearchBot are not disallowed | Engineering | 1 day | A |
| 4 | Audit Cloudflare/WAF settings — deadline 15 September 2026 | Engineering | 1 day | A |
| 5 | View-source and curl-with-bot-user-agent test: is your key content in the initial HTML? | Engineering | 2 days | A |
| 6 | Pull AI-bot server logs; quantify 404 and redirect waste | Engineering | 2 days | A |
| 7 | Enable GSC generative AI reports and Bing AI Performance report | SEO | 1 hour | A |
| 8 | Verify GA4 AI Assistant channel; document the Direct-attribution gap | Analytics | 1 day | A |
| 9 | Harvest real customer language from sales calls, support, site search | Content/Sales | 1 week | B+ |
| 10 | Run a baseline accuracy audit: read 50 AI answers about your brand | Marketing | 2 days | A |
| 11 | Brief legal on the FTC review rule, Operation AI Comply, EU AI Act Article 50 | Legal | 1 day | A |
| 12 | Name a single accountable owner and schedule the quarterly forum | CMO | 1 day | — |
End-of-month deliverable: a one-page diagnostic — number of entity contradictions, percentage of key content server-rendered, AI crawler 404 rate, baseline accuracy findings, and whether any platform is currently blocked by misconfiguration.
That last item is the one that most often justifies the whole exercise. Blocking OAI-SearchBot while intending to block GPTBot is a common, invisible, entirely self-inflicted wound.
Days 31–60: Fix and build the baseline
| # | Action | Owner | Effort | Grade |
|---|---|---|---|---|
| 13 | Fix entity contradictions: LinkedIn, Crunchbase, directories, press boilerplate | Brand/Comms | 3 weeks | A− |
| 14 | Add a plain-text identity statement to homepage or About page | Brand | 1 day | A− |
| 15 | Implement Organization schema with complete sameAs | Engineering | 1 day | A− |
| 16 | Create or complete a referenced Wikidata item; record the Q-ID | SEO | 1 day | A− |
| 17 | Move client-side-only primary content to server-side rendering | Engineering | 2–4 weeks | A |
| 18 | Fix the top AI-crawler 404s and redirect chains | Engineering | 3 days | A |
| 19 | Build the 120–200 prompt set from the Day 1–30 harvest | Content | 1 week | B+ |
| 20 | Buy one Family B and one Family A measurement tool; ask both for run counts | SEO | 2 weeks | B |
| 21 | Map the third-party roundups and review pages cited for your top 10 commercial queries | PR/SEO | 1 week | B+ |
| 22 | Correct factual errors about you in those third-party sources | PR | 2 weeks | B+ |
| 23 | Audit product feed completeness; populate Product Highlight and Product Detail | Ecommerce | 2 weeks | A |
| 24 | Review bot mitigation against agentic browser user agents | Security/Eng | 1 day | B+ |
End-of-month deliverable: a measurement baseline with 14-day rolling mention rate and citation rate reported separately, plus a corrected entity footprint.
Do not report a trend yet. You do not have one. Reporting a trend from three weeks of data is exactly the error Chapter 10 warns about, and doing it once establishes a precedent you will spend a year unwinding.
Days 61–90: Build the compounding assets
| # | Action | Owner | Effort | Grade |
|---|---|---|---|---|
| 25 | Publish Tier 1 constrained-recommendation pages for the top 20 attribute intersections | Content | 6 weeks | B |
| 26 | Publish prices and specifications wherever commercially possible | Content/Product | 2 weeks | B+ |
| 27 | Commission the first original research study | Marketing | starts now | B |
| 28 | Launch a digital PR program aimed at the outlets that actually get cited | PR | ongoing | B+ |
| 29 | Establish a named, disclosed community presence on the two forums where your category is discussed | Community | ongoing | B |
| 30 | Launch a sentiment-neutral review solicitation program | CX | ongoing | B |
| 31 | Fund a YouTube pilot with real substance | Content | ongoing | B− |
| 32 | Migrate FAQ content into feed Question and Answer attributes; attach Document Links | Ecommerce | 2 weeks | A |
| 33 | Publish /.well-known/ucp if selling where Universal Cart is live | Engineering | 1 sprint | A− |
| 34 | Run the first quarterly agent-completability test | Ecommerce | 1 day | B+ |
| 35 | Institute a content refresh cycle; recency is a retrieval factor | Content | ongoing | B |
| 36 | First quarterly cross-functional forum: review inputs and outputs together | CMO | half day | — |
End-of-quarter deliverable: a report with four output metrics on rolling averages, five input metrics you control, an accuracy finding, and an explicit statement of what you still cannot measure.
That last section — what you cannot measure — is the most credible thing in the document. It is also the thing that protects you when someone asks why the numbers moved and the honest answer is that they did not move, they varied.
What not to do in the first 90 days
Do not restructure the content library for chunkability. Grade C.
Do not commission a schema program aimed at AI citations. Grade C−.
Do not pay anyone for llms.txt. Grade D.
Do not block Google-Extended expecting it to affect AI Overviews. Grade D.
Do not pull the Search Console AI opt-out.
Do not seed reviews, forum posts or testimonials. Grade F for the practice, and the corresponding control — never seeding — is graded A. Statutory maximum $53,088 per knowing violation.
Do not report a trend before you have 14 days of properly sampled data.
Do not fund an “agentic commerce transformation.” Fund a feed and a manifest.
The one-sentence version
If you do nothing else in ninety days: check your crawler configuration, make your identity consistent everywhere it appears, put your real content in the initial HTML, and start measuring properly. Those four things are graded A, cost very little, and are undone in most organizations right now.
Everything else in this book is optimization on top of them.
BACK MATTER
Conclusion: What This Is Actually About
The temptation with a shift this large is to treat it as a new channel to be conquered — a new acronym, a new tool category, a new line item, a new agency. That reading is wrong, and it is expensive.
The evidence, read honestly, says something quieter and more demanding.
Most of what is sold as AEO is either traditional excellence renamed, or unevidenced speculation. Getting retrieved dominates everything downstream: roughly 70% of AI Overview citations come from domains already on page one, pages ranking first get cited at 3.5 times the rate of pages beyond position twenty, and in the one peer-reviewed head-to-head, traditional SEO outperformed purpose-built conversational tactics. The content-cosmetic layer that most of this industry sells — restructuring, schema, chunk formatting — measures at or near zero in the best controlled experiments available.
What is genuinely new is smaller and harder. Four things:
The entity layer, because a machine cannot apply a credibility signal to something it cannot identify, and because most prompts never trigger a search at all — meaning what the model already believes about you is the whole game for the majority of queries.
The corroboration layer, because grounding confidence is derived partly from how many independent documents verify a claim. Third-party agreement is not reputation management in the soft sense. It is an input to a scoring function.
The specificity layer, because query fan-out means you compete on questions nobody typed, and because 65% to 85% of prompts do not match any keyword in the largest keyword database in the industry.
The agent layer, because a brand an assistant cannot read, price, or transact with is a brand it will route around.
And what is required to manage all of it is a tolerance for uncertainty that most marketing organizations are not built for. You will be asked to invest in a channel where impressions are visible but clicks are not, where the same measurement run twice disagrees with itself, where the dominant effect is influence you cannot observe, and where the best available research contradicts the most-cited research. The organizations that handle this well will be the ones that report honestly — rolling averages, separated metrics, explicit statements of what is not known — rather than the ones that produce the most confident dashboard.
There is a version of this book that would have been easier to write and more fun to read. It would have said the click is dead, the rules are rewritten, and here are eleven tactics to win. Some of those tactics would have been the ones that measure at zero.
The actual instruction is less dramatic and more useful: be a legible entity, be reachable and readable, be corroborated by people who do not work for you, be specific enough to answer the real question, and be usable by a machine acting on a customer’s behalf.
The brands that do those five things will be the answer. Not because they optimized for it, but because when a system searches for a credible, consistent, specific, corroborated source, that is what it finds.
Appendix A: The AEO Audit
A single-pass diagnostic. Score each item pass, partial, or fail. Anything graded A that scores fail is an emergency; anything graded C or below that scores fail is a low priority — except rows phrased as prohibitions (“Nobody has…”, “No mass-generated…”), where a fail means you are doing something the book says not to do, and the grade describes the tactic, not the urgency of stopping.
Entity
| Check | Grade |
|---|---|
| A canonical facts document exists and has positioning sign-off | A− |
| Homepage or About page states identity in plain indexable text | A− |
| Organization schema present with complete sameAs | A− |
| Wikidata item exists, is referenced, and its Q-ID is in sameAs | A− |
| LinkedIn, Crunchbase and directory descriptions match the canonical statement | A− |
| Press boilerplate in the last ten releases matches current positioning | A− |
| Named experts have author pages, consistent bylines and author markup | B |
| Knowledge Panel claimed, if one exists | B |
| Quarterly consistency review has a named owner | A− |
Retrievability
| Check | Grade |
|---|---|
| robots.txt does not disallow OAI-SearchBot, PerplexityBot or Claude-SearchBot | A |
| Nobody has blocked GPTBot believing it controls citation | A |
| Nobody has blocked Google-Extended believing it controls AI Overviews | D (prohibition) |
| Cloudflare/WAF settings audited against the 15 September 2026 defaults | A |
| Key content appears in view-source and in a curl fetch using each bot’s user agent | A |
| AI-crawler 404 rate measured and top offenders fixed | A |
| No noindex or nosnippet on pages intended for AI answers | A |
| Bot mitigation reviewed against agentic browser user agents | B+ |
| Search Console AI opt-out not enabled | A |
Corroboration
| Check | Grade |
|---|---|
| Third-party roundups and review pages cited for top 10 commercial queries are mapped | B+ |
| Factual errors about you in those sources have been corrected | B+ |
| A digital PR program targets the outlets that actually get cited | B+ |
| At least one original research study is published or commissioned annually | B |
| A named, disclosed community presence exists on the two key forums | B |
| A sentiment-neutral review solicitation program is running | B |
| No reviews, forum posts or testimonials have ever been seeded or incentivized by sentiment | A (F if failed) |
| FTC review rule is in agency contracts | A |
| A YouTube program exists with real substance | B− |
Specificity
| Check | Grade |
|---|---|
| A prompt inventory built from real customer language exists | B+ |
| Attribute intersections are mapped, not just head terms | B |
| Prompts scored for winnability: citations, rank, honest information gain | B+ |
| Prices and specifications published wherever commercially possible | B+ |
| Tier 1 constrained-recommendation pages exist for top intersections | B |
| Fair comparison content exists, including against competitors | B |
| Content is dated and on a defined refresh cycle | B |
| No mass-generated permutation pages | C− (prohibition) |
Transactability
| Check | Grade |
|---|---|
| Product feed complete and refreshed on a defined cadence | A |
| Product Highlight and Product Detail populated across catalogue | A |
| Q&A attributes populated in the feed | A |
| Manuals and spec sheets attached via Document Link | A |
| /.well-known/ucp published, where Universal Cart is live | A− |
| Price, availability and fulfilment terms in initial HTML | A |
| Quarterly agent-completability test run on core journeys | B+ |
Measurement
| Check | Grade |
|---|---|
| GSC generative AI reports and Bing AI Performance report enabled | A |
| GA4 AI Assistant channel verified, with the Direct gap documented | A |
| AI-bot log analysis running | A |
| Prompt set of 120–200, versioned and frozen quarterly | B+ |
| One Family A and one Family B tool in place | B |
| Only two-to-four-week rolling averages reported | A |
| Mention rate and citation rate reported separately | A |
| Monthly accuracy audit of answers about the brand | A |
| Controllable inputs reported alongside outputs | B+ |
| Survey questions capturing AI-assisted research added | B |
Appendix B: Prompt Set Template
Structure
Target 120–200 prompts. Distribution:
Fields to record per prompt
| Segment | Share | Count at 150 |
|---|---|---|
| Constrained recommendation | 40% | 60 |
| Comparison | 20% | 30 |
| Brand-specific | 15% | 23 |
| Implementation and troubleshooting | 15% | 22 |
| Category education | 10% | 15 |
For each core buying question, enumerate constraints and generate intersections:
Run each prompt at least seven times per day; report only rolling averages.
Attribute matrix worksheet
Cover the full attribute matrix, not just the head terms.
Freeze for a quarter. Version any change and treat the series as broken at that point.
Include prompts you currently lose. A set built from strengths teaches nothing.
Include your top three competitors’ strongest positions.
Construction rules
Source from harvested customer language, not from a brainstorm.
| Field | Notes |
|---|---|
| Prompt text | Verbatim, in customer language |
| Segment | From the table above |
| Source | Sales call, support ticket, site search, forum, competitor position |
| Commercial value | High / medium / low |
| Current rank for the closest search equivalent | From Search Console |
| Baseline cited domains | Recorded across at least seven runs |
| Baseline mention (yes/no) | Averaged across runs |
| Information gain we hold | Honest answer, or “none” |
| Winnable | Yes / no / not yet |
| Dimension | Your values |
|---|---|
| Company size / household type | |
| Industry or use context | |
| Integration or compatibility requirements | |
| Budget band | |
| Deployment or delivery model | |
| Compliance or regulatory regime | |
| Buyer sophistication | |
| Specific job to be done |
Three to five dimensions with three to five values each generates more intersections than you can serve. Rank by commercial value and honest winnability, then build the top twenty.
Appendix C: The Evidence Table
Every tactic in this book, graded. Where two studies conflict, both are named.
| Tactic | Best evidence | Grade |
|---|---|---|
| Server-side render primary content | Vercel, Dec 2024 (1.3B requests); searchVIU, Nov 2025 (23 crawlers, 69% cannot execute JS) | A |
| Fix AI-crawler 404s and redirect chains | Vercel: ~35% of ChatGPT and Claude fetches hit 404s | A |
| Correct crawler access control by user agent | First-party docs from OpenAI, Anthropic, Perplexity, Google | A |
| Product feed completeness and the eight Merchant Center attributes | Google first-party specification, stated AI-surface purpose | A |
| Do not seed or incentivize reviews | FTC rule, effective October 2024; live enforcement 2026 | A |
| Do not pull the Search Console AI opt-out | No click or CTR data to decide on; comparative ranking system | A |
| Report rolling averages only | Schulte et al.: Jaccard 0.34–0.42 run-to-run | A |
| Entity consistency and Wikidata | Documented grounding mechanism; strong correlational support | A− |
| /.well-known/ucp manifest | Google first-party spec; Universal Cart live in the US | A− |
| Earned third-party mentions | Ahrefs: r ≈ 0.66, three times backlinks; Muck Rack: 82% of citations earned | B+ |
| Digital PR aimed at cited outlets | Muck Rack; AirOps third-party listicle dominance (80.9%) | B+ |
| Agent-completability testing | HUMAN Security agentic traffic composition | B+ |
| Organizing coverage around real question-form queries | Xu et al.: 6.8x activation multiplier on query form (not page formatting) | B |
| Original research and data | GEO paper (cite sources +27%); earned-media citation patterns | B |
| Publishing prices and specifications | Sprinklr: odds ratio >10,000 for price presence, in a synthetic testbed | B+ |
| Content recency and refresh cycles | Seer: 65% of AI bot hits target content under one year old | B |
| Passage self-containment | Documented chunking mechanism; direct effect unmeasured | B− |
| Statistics and quotations in content | GEO paper (+31%, +41%), simulated engine; contradicted by C-SEO Bench | B− |
| YouTube presence | Ahrefs: strongest correlate (r ≈ 0.74); BrightEdge 20% citation share | B− |
| Structural rewriting for chunkability | Sprinklr: negligible. C-SEO Bench: sometimes harmful | C |
| JSON-LD schema as a citation lever | Ahrefs DiD (n=1,885): −4.6% to +2.4%; Google: no special markup needed | C− |
| Mass-generated permutation pages | Ahrefs: content volume r ≈ 0.17–0.19 | C− |
| llms.txt | Zero frontier-crawler fetches in 900-domain, 7-month log study | D |
| Blocking Google-Extended to control AI Overviews | Google docs: does not affect Search inclusion or AI features | D |
| Keyword stuffing | GEO paper: −8%, worst method tested | D |
| Manipulating AI recommendations | ETH Zürich: works, then degrades for everyone; FTC exposure | F |
Appendix D: Glossary
AEO / GEO / GAIO / LLMO — Agentic, Generative, or Large Language Model Engine Optimization. Used interchangeably in practice. This book uses AEO, following the source paper, because “agentic” captures the direction of travel toward systems that act rather than only answer.
Algorithmic Trinity — Jason Barnard’s framing of answer engines as three interdependent systems: the language model (synthesis), the search index (retrieval), and the knowledge graph (validation).
Chunk — a segment of a document, typically a few hundred tokens, independently embedded and indexed. The actual unit of retrieval.
Citation rate — the share of tracked prompts in which your domain appears as a source. Distinct from mention rate.
Entity Home — the single authoritative brand-owned property where systems find definitive facts about your identity. Usually the homepage or About page.
Expected rank — E[E[rank | user]]. The average of personalized expected rankings across the user population. Replaces static rank as the meaningful visibility metric under personalization.
Ghost citation — an appearance where your site is cited as a source but your brand is never named in the answer. 61.7% of appearances in the Semrush study.
Grounding — the process of checking generated claims against retrieved sources. Produces support scores and can suppress low-confidence output.
Information gain — the marginal value a passage offers over the passage it is compared against. The one content principle that survives the evidence.
Mention rate — the share of tracked prompts in which your brand is named in the answer. The recommendation metric.
NEEATT — Notability, Experience, Expertise, Authoritativeness, Trustworthiness, Transparency. Barnard’s extension of Google’s E-E-A-T for machine evaluators.
Pairwise reranking — ranking by repeated head-to-head comparison rather than absolute scoring. Documented in research and patented by Google; inferred, not confirmed, in production.
Query fan-out — decomposing a single prompt into multiple synthetic sub-queries executed in parallel. Officially confirmed by Google; the number of sub-queries is not.
RAG — retrieval-augmented generation. Retrieve relevant passages, put them in the model’s context, generate an answer grounded in them.
Retrieval threshold — the score below which a system answers from its own knowledge without searching. Google’s documented default is 0.3.
Share of voice — frequency and prominence of brand appearance in AI answers relative to competitors. Widely sold, rarely sampled adequately.
UCP / ACP / MCP / A2A / AP2 — Universal Commerce Protocol (Google), Agentic Commerce Protocol (OpenAI and Stripe), Model Context Protocol (Anthropic-originated), Agent2Agent (Linux Foundation), Agent Payments Protocol (Google).
Appendix E: Sources and Further Reading
The foundation
Bliey, Miles, and Keira Chatwin. Answer Engine Optimization: How Agentic AI Reshapes SEO. Stanford GSBGEN 390, July 2026. SSRN.
The research that matters most
Aggarwal, Pranjal, et al. “GEO: Generative Engine Optimization.” Proceedings of the 30th ACM SIGKDD Conference, 2024, pp. 5–16. The founding study. Read Table 1 yourself.
Puerto, Haritz, et al. “C-SEO Bench: Does Conversational SEO Work?” NeurIPS 2025, Datasets and Benchmarks Track. The peer-reviewed challenge to the above. Read both together or neither.
Vishwakarma, Rahul, et al. “What Gets Cited: Competitive GEO in AI Answer Engines.” SIGIR 2026. 252,000 controlled trials.
Xu, Haofei, Umar Iqbal, and Jacob M. Montgomery. “Measuring Google AI Overviews.” Washington University in St. Louis, arXiv, May 2026. The largest audit of real AI Overviews.
Schulte, Julius, Malte Bleeker, and Philipp Kaufmann. “Don’t Measure Once: Measuring Visibility in AI Search.” University of St. Gallen, arXiv, April 2026. The measurement study nobody selling dashboards wants you to read.
Qin, Zhen, et al. “Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting.” Findings of NAACL 2024, pp. 1504–1518.
Nestaas, Fredrik, Edoardo Debenedetti, and Florian Tramèr. “Adversarial Search Engine Optimization for Large Language Models.” ETH Zürich, arXiv 2406.18382.
Primary platform documentation
Google Search Central, AI features documentation — the single most useful page in this field, and the one that contradicts the most vendor decks.
Google Cloud, Vertex AI Search parsing and chunking documentation; Check Grounding API documentation.
Google patents: US11769017B1 (Generative summaries for search results, granted); US20240289407A1 (Search with stateful chat — the fan-out one); US20250124067A1 (Pairwise ranking prompting).
OpenAI bots documentation and commerce specifications; Anthropic crawler documentation; Perplexity crawler documentation.
Google Merchant Center attribute documentation; Universal Commerce Protocol developer documentation.
Cloudflare, “Content Independence Day: AI options,” 1 July 2026.
Measurement and industry research
Ahrefs: brand visibility correlations (~75,000 brands); schema difference-in-differences study; most-cited domains by platform; AI Overview CTR impact. Vendor research with unusually transparent methodology.
Profound: AI platform citation patterns; “How ChatGPT sources the web.” Note the top-ten-versus-all-citations distinction.
Similarweb: AI Search Stats 2026; most-cited domains; generative AI landscape.
Semrush: ChatGPT search insights (clickstream); the ghost citations study with Kevin Indig.
SparkToro: zero-click search analysis, June 2026.
Pew Research Center: “Google users are less likely to click on links when an AI summary appears,” July 2025. The one measurement here with no commercial interest.
Muck Rack: GenerativePulse 2025. Largest earned-media citation dataset.
Vercel: “The rise of the AI crawler,” December 2024. searchVIU: schema retrieval tests and AI crawler JavaScript analysis.
Legal and regulatory
FTC Trade Regulation Rule on the Use of Consumer Reviews and Testimonials, effective 21 October 2024, and the revised Endorsement Guides, June 2023.
EU AI Act Article 50 transparency obligations, in force 2 August 2026.
Moffatt v. Air Canada, 2024 BCCRT 149. Walters v. OpenAI, Georgia, May 2025. Starbuck v. Google, filed October 2025. Regional Court of Munich, case 26 O 869/26, May 2026.
Wikipedia:Conflict of interest, and Wikidata:Notability. Read both before anyone on your team touches either site.
Practitioner analysis worth following
Michael King (iPullRank) on AI Mode architecture and relevance engineering. Olaf Kopp on patents, query fan-out and brand context optimization. Jason Barnard (Kalicube) on the Algorithmic Trinity, NEEATT and the Entity Home. Kevin Indig (Growth Memo) on citation concentration and platform data. Chris Long on digital authority.
Treat all practitioner analysis as informed inference from patents and observation. It is valuable, it is frequently right, and it is not documentation.
A closing note on the figures in this book. Every quantitative claim was verified against primary sources in August 2026. Several will be wrong by the time you read this, because platform behaviour, market share and regulation in this field move faster than publishing does. The evidence grades will age better than the numbers. When a figure here conflicts with something newer, trust the newer one — and check its denominator.
Put the book to work
The levers in this guide are the same ones we run for clients. Start with generative engine optimization, win the direct answer with answer engine optimization, fix retrieval with technical SEO, earn consensus through digital PR, and keep score with AI visibility monitoring. Our methodology maps chapter by chapter.
Read next on visibility.partners
- Generative Engine Optimization servicesChapters 5–8 delivered as a retained program.
- Answer Engine Optimization servicesWinning the direct answer and the sub-query.
- Technical SEO servicesThe retrievability work from Chapter 6.
- Digital PR servicesCorroboration and citation sources from Chapter 7.
- AI visibility auditAppendix A, run for your brand across five engines.
- AI visibility monitoringAppendix B prompt sets, re-run monthly.
- Our methodologyHow the book's 90-day plan maps to an engagement.
- What is answer engine optimization?A short primer on the ideas in Chapters 2 and 8.
Primary sources and further reading
Every claim in this book is grounded in vendor documentation, published research, or our own measurement. These are the primary sources worth reading directly.
- OpenAI platform documentation — How ChatGPT retrieval, browsing and tool use are documented by OpenAI.
- OpenAI: search in ChatGPT — OpenAI's own description of search behaviour and publisher citation.
- Google Search Central: AI features and your website — Google's guidance on AI Overviews, AI Mode and eligibility.
- Google: creating helpful, reliable, people-first content — The quality guidance underpinning E-E-A-T signals.
- Google structured data reference — Supported schema types for rich results and machine parsing.
- Schema.org vocabulary — The entity vocabulary used throughout Chapters 5 and 6.
- Perplexity developer documentation — How Perplexity retrieves, ranks and attributes sources.
- Google Gemini API documentation — Grounding, search retrieval and citation behaviour in Gemini.
- Anthropic documentation — Claude's retrieval, citations and tool-use model behaviour.
- Bing webmaster guidelines — Indexing and quality guidance behind Copilot's web answers.
- GEO: Generative Engine Optimization (Aggarwal et al., 2023) — The peer-reviewed paper that named generative engine optimization.
- Retrieval-Augmented Generation (Lewis et al., 2020) — The original RAG paper behind the retrieval pipeline in Chapter 2.
- web.dev: Core Web Vitals — Performance thresholds referenced in the retrievability audit.
- The robots.txt specification — Baseline for AI crawler access decisions in Chapter 6.
Ranking is no longer enough
You need to be cited, mentioned, and recommended.
Get a free AI Visibility Report — see exactly where your brand appears across ChatGPT, Google AI, Gemini, Perplexity, and Copilot, and where competitors are winning instead.