Resources · Complete guide
Becoming the Answer: How Brands Get Chosen Inside AI Search
The full manuscript by Jeremy Osborn — published free. Fourteen chapters and five appendices on the retrieval pipeline, the five levers of AI visibility, what the evidence actually supports, and how to run the program.
21 chapters · ~19k words · August 2026

About this guide
Becoming the Answer
How brands get chosen when AI does the choosing
Jeremy Osborn
August 2026
PART ONE — THE NEW FRONT DOOR
Chapter 1 — The Question That Never Reaches You
A woman in Ohio needs a new payroll system for her fourteen-person dental practice. Three years ago she would have typed “best payroll software small business” into Google, scrolled past four ads, opened six tabs, and spent an evening comparing pricing pages.
Tonight she types something different:
“What payroll system should a 14-person dental practice use if we have two part-time hygienists, need to handle tips, and already use QuickBooks?”
Fifteen seconds later she has three names, a reason for each, and a note about which one handles tipped employees badly. She opens one tab. She books a demo.
Two of the three vendors in that answer have no idea it happened. The third has no idea either, but it got the demo.
Nothing in that vendor’s analytics will show what occurred. There is no keyword. There is no impression they can price. There is a single session that arrives already convinced, from a referrer most attribution models will file under Direct.
This is the shape of the problem. Not that customers stopped searching — they search more than ever. The shape is that the deciding moment moved somewhere you cannot see.
The arrangement that just ended
For twenty-seven years, search worked on an unwritten deal.
Search engines organized the world’s information. Companies produced it. People traveled between the two through a list of blue links. Everything the marketing profession built on top of that — rank tracking, click-through optimization, attribution modeling, the whole machinery of performance marketing — rested on one assumption: the search engine was a doorway, and value was created when someone walked through it.
That assumption is failing. Not because people stopped using Google. Because the doorway learned to answer.
The scale of the shift is not in dispute. ChatGPT was approaching a billion weekly users by mid-2026. Google’s Gemini app crossed a billion monthly users in August of the same year. Across all the AI assistants, Similarweb counted 9.5 billion monthly visits from 655 million people, up about 70 percent year over year.
But adoption is the least interesting part of the story, and it is the part every conference keynote stops at. Three findings underneath it matter far more.
Nobody switched
The first is that people did not leave Google for ChatGPT. They added ChatGPT to Google.
Similarweb’s tracking shows 95 percent of ChatGPT users also still use Google — a number that has not moved in a year. Search did not lose a war. It gained a neighbor.
Which means any strategy built on “AI is replacing search, so move the budget” is built on something that has not happened and shows no sign of happening soon.
The assistants barely send traffic
The second finding is more uncomfortable, and it is the one that reorganizes budgets.
AI platforms send almost nothing. Ahrefs, measuring roughly 82,000 websites, found AI referrals averaging a quarter of one percent of total traffic. Microsoft’s own study of 1,277 domains put it at “less than one percent.” Scrunch, watching millions of visits across news publishers, found 1.1 percent carrying an AI referrer, against roughly nine percent from traditional search and seventy-five percent direct.
The growth rates are enormous — Ahrefs measured a nearly tenfold year-over-year increase, Adobe measured 393 percent growth in AI-referred retail traffic in a single quarter — but they are enormous growth from a rounding error.
If your plan for AI search is a referral acquisition program, you are building a channel that currently delivers one visitor in four hundred.
Meanwhile, the traffic you already had is disappearing
The third finding is where the damage actually is.
In the first four months of 2026, 68 percent of US Google searches ended without a single click. That is up from 60 percent in 2024 and 49 percent in 2019.
Pew Research — the one organization measuring this with no product to sell — watched real people browse across 68,879 searches. When an AI summary appeared at the top of the results, people clicked a link 8 percent of the time. When it didn’t, they clicked 15 percent of the time. They clicked a link inside the AI summary 1 percent of the time.
Ahrefs compared 300,000 keywords across two years of Search Console data and found that first-position click-through rate on AI Overview keywords fell 58 percent against what it should have been. Publishers report roughly a third less referral traffic than a year ago.
Line those three findings up and the real picture appears, and it is not the one on the conference slides:
The chatbots are not taking your traffic. Google’s own AI features are absorbing the clicks that used to leave. The chatbots are quietly deciding who gets recommended, and sending almost nobody to prove it.
The good news, with a shelf life
The traffic that does arrive from AI platforms behaves unusually well.
Adobe, working with large-scale retail data, found AI-referred visitors converting 42 percent better than everyone else, spending 48 percent more time on site, and producing 37 percent more revenue per visit. The Washington Post’s revenue chief reported AI-platform visitors subscribing at four to five times the rate of search visitors.
There is a catch, and it is worth knowing before you build a business case on it. Twelve months earlier, Adobe’s own data showed AI-referred traffic converting at roughly half the rate of everything else. The advantage is recent. It reflects a moment when the people using AI assistants to research purchases skew toward high-intent, high-income, technically confident buyers. As that population broadens, expect the advantage to compress.
Ahrefs also found AI visitors bouncing slightly more than search visitors and viewing fewer pages. These are not superhumans. They are people who did their research somewhere else and arrived with the decision mostly made.
Which is the actual insight. The click is not more valuable because the visitor is better. The click is more valuable because the persuasion already happened, out of view, in a conversation you were not part of.
What this book is for
If the deciding moment happens inside a system you cannot observe, the only lever you have is what that system finds when it goes looking.
That is a smaller and more concrete problem than it sounds. It comes down to five questions, and the rest of this book answers them:
Can the machine tell who you are? Can it reach and read what you published? Does anyone independent back up what you claim? Do you answer the exact question, not the general one? Can an assistant actually use you on a customer’s behalf?
Get those right and you show up. Get them wrong and you can outspend everyone in your category and remain invisible to the woman in Ohio, who will never know you existed, and who will never appear in your funnel as a lost opportunity, because she was never in it.
Chapter 2 — Where the Traffic Went
Before anything else, you need a clear-eyed account of what is actually happening to your numbers, because the wrong diagnosis leads to expensive wrong medicine.
Three things are happening at once, and they are frequently blamed on each other.
Thing one: Google is answering more questions itself
AI Overviews now appear on somewhere between one in five and two in five US searches, depending on who is measuring and what mix of queries they sampled. The spread is wide because query type drives everything.
A rigorous audit out of Washington University in St. Louis captures this well. Across 55,393 trending queries over forty days, AI Overviews appeared on 13.7 percent overall. But on queries phrased as questions, they appeared 64.7 percent of the time — nearly seven times as often. By category the range ran from 3.5 percent in beauty and fashion to 46 percent in hobbies and leisure, with politics and legal queries suppressed down near 8 percent.
So “how much of my category is affected” is a question with a real answer, and it is probably not the number in the headline you read. Run your own top hundred queries and count.
Thing two: your analytics is undercounting AI traffic
The AI referral traffic you can see is a floor, not a total.
Google Analytics added an AI Assistant channel group in 2026, which sets the medium to ai-assistant when it recognizes the referrer. Three problems: traffic from in-app browsers and mobile apps often arrives with no referrer at all and lands in Direct; Google has not published the list of assistants it recognizes; and it only captures referrals, not the far larger effect of answers where nobody clicks.
One nine-month first-party study tracking 51,200 AI Overview click events found 22 percent of that traffic misattributed to Direct rather than organic search. Whatever your dashboard says, the real number is higher.
Thing three: the assistants themselves are changing fast
The competitive picture is unstable in a way that should make you wary of platform-specific tactics.
Between mid-2025 and mid-2026, ChatGPT’s share of generative AI web traffic fell from roughly 76 percent to about 53 percent. Gemini rose from under 9 percent to around 27. Claude went from about 2 percent to 9. Perplexity and Copilot sit in low single digits.
A twenty-point swing in twelve months is not a market you tune a strategy to. It is a market you build platform-independent assets for.
The question that actually gets asked
Here is a finding that should change how content teams work, and it is not widely known.
Semrush, analyzing over a billion lines of clickstream data from a 200-million-user panel, found that between 65 and 85 percent of ChatGPT prompts do not match any keyword in a database of 27 billion keywords.
Read that again if your content plan starts with a keyword tool.
The majority of demand you are trying to serve is invisible to your primary instrument, because people phrase questions to an assistant the way they’d phrase them to a knowledgeable friend. Queries in Google’s AI Mode run about three times longer than classic search queries. People type “best CRM software.” They ask “what CRM should a twelve-person agency use if we bill on retainers and need QuickBooks to sync.”
Keyword research still tells you what people type into search boxes. It no longer tells you what they ask.
The gap you are being asked to manage
Put all this together and you arrive at the honest position, which most vendors will not lead with:
Google Search Console will show you AI impressions but not clicks, click-through rate, queries, or position. Bing’s tools will show you citations but not traffic value. Analytics will catch some referrals and lose the rest to Direct. And the dominant effect — being recommended in an answer nobody clicks — leaves no trace anywhere.
You are being asked to manage a channel where the cost is measurable and the benefit is not.
That is uncomfortable, and it is also the actual job. Chapter 11 lays out how to do it without lying to yourself or your board. But the discomfort is worth naming early, because every bad decision in this field starts with someone pretending the gap isn’t there.
What not to conclude
Two wrong turns are common at this point.
“AI search is hype, the traffic is a rounding error, ignore it.” The traffic is a rounding error. The influence is not. Sixty-eight percent of searches ending without a click means the research is happening somewhere; it just isn’t happening on your site anymore. Waiting for the referral numbers to justify the investment means waiting until the positions are taken.
“AI search changes everything, rip up the playbook.” It doesn’t and you shouldn’t. As Chapter 4 shows, most of what determines whether you appear in an AI answer is whether you were findable in the first place. The playbook gets a new chapter, not a new cover.
The right posture is in between and slightly boring: treat AI visibility as a consequence of being a well-defined, well-corroborated, technically accessible business — and fix the parts of that you have been neglecting.
Which requires knowing what the machine actually does. That’s next.
Chapter 3 — What Happens in the Two Seconds Before an Answer

You don’t need to be an engineer to run this well. You do need an accurate picture of what happens between a question and an answer, because nearly every bad tactic in this field comes from an inaccurate one.
Here is the sequence, in plain terms.
Step one: it decides whether to look anything up
Before anything else, the system decides whether to search the web at all.
This is documented rather than inferred. Google’s grounding system assigns each prompt a score between 0 and 1 estimating whether searching would improve the answer, with a default threshold of 0.3. Below that, the model answers from what it already knows. No search, no sources, no citations.
The observed rates are striking. Profound, analyzing roughly 730,000 real ChatGPT conversations, found only 18 percent triggered a web search. Semrush’s clickstream data put web search enabled on 34.5 percent of ChatGPT queries — down from 46 percent a year earlier.
Sit with that for a second.
Somewhere between two-thirds and four-fifths of questions are answered from memory. For those questions, your website is irrelevant. Not underperforming. Irrelevant. What matters is whether the model already knows your company exists, what category it belongs to, and what it’s good at.
That single fact is the reason Chapter 5 comes before Chapter 8. Being known is upstream of being read.
Step two: it breaks the question apart
When the system does search, it usually doesn’t run one search. It splits the question into several sub-questions and runs them in parallel. Google calls this query fan-out, and describes it openly: one engineering director put it as “doing a dozen searches for you in the time it takes to do one.”
Nobody outside Google knows how many sub-queries, and anyone who tells you a specific number is guessing. But the consequence holds regardless.
Our dental-practice question fans out into something like: payroll systems for small businesses, payroll for medical and dental practices, handling tipped employees, part-time employee payroll, QuickBooks payroll integrations, pricing for teams under twenty.
You might win four of those and lose on tipped employees, and lose the answer.
You are competing on questions nobody typed and no keyword tool will ever show you.
Step three: it retrieves passages, not pages
Each sub-question returns candidate passages. Not pages — passages.
Documents get chopped into chunks of a few hundred tokens each, indexed independently. Google’s search chunker caps at 500 tokens. Microsoft recommends starting at 512. Google’s own retrieval product defaults to 1,024 with overlap.
So your 3,000-word guide does not compete as a guide. It competes as roughly eight separate fragments, most of which will never be seen together.
This has a practical edge. A section that opens “as we discussed above” is ambiguous when it arrives alone. The systems do some work to mitigate this — headings can be attached to chunks, and neighboring chunks can be pulled in alongside a match — so the situation is less dire than some presentations suggest. But if a section can’t stand up by itself, don’t count on it.
Step four: it compares passages against each other
Surviving passages get re-ranked before they reach the model, and there is good evidence the comparison is head-to-head rather than absolute.
Google Research published work showing that asking a model to compare two documents at a time — a much simpler task than scoring one in isolation — lets a modest open model match GPT-4’s re-ranking quality. Google then patented the method.
That is a methods paper, not a description of what runs in production, and it’s worth being precise about the difference. But it points at something that matches everything else we observe: your passage doesn’t need to be good. It needs to beat the specific passage it’s compared against.
Which reframes the whole content question. A verifiable number, a named source, a real price, a dated measurement, a specific example — these give a comparison something to prefer. A well-written paraphrase of common knowledge gives it nothing.
Practitioners call the difference information gain. It is the most transferable idea in this entire field.
Step five: it checks the claims
Passages that survive get checked against the answer being drafted. This is documented in detail.
Google’s grounding check produces a support score from 0 to 1 for how well an answer is grounded in the retrieved facts, with a default citation confidence threshold of 0.6. Claims get mapped to source chunks at roughly sentence granularity.
Google’s granted patent on generative summaries goes further, and this is the part with strategic consequences. Confidence for each portion of the summary is derived from three things: the model’s own confidence, the trustworthiness of the supporting documents, and how many documents verify that portion. Low-confidence summaries can be suppressed entirely.
Read that last mechanism again.
How many independent sources agree with a claim is an input to whether the claim survives.
A fact stated only on your website is a single-source claim. The same fact appearing on your site, in a trade publication, in a directory listing, and in a structured knowledge base is a multi-source claim. The architecture prefers the second, mechanically, not as a matter of taste.
That is the entire argument of Chapter 7, and it is not a marketing metaphor. It is a scoring function.
Step six: it writes, differently every time
Validated passages go into the model’s context with source identifiers attached, which is how specific sentences link back to specific documents.
The output is probabilistic. Run the same prompt twice and you get different answers, and the variance is larger than most people assume. Researchers at the University of St. Gallen found that identical prompts, run on the same day, produced source overlap of only 0.34 to 0.42 and brand-mention overlap of 0.45 to 0.59. The instability held inside 24-hour windows, ruling out news cycles.
Their recommendation: at least seven runs of a prompt per day before you believe a number.
Almost no commercial dashboard does that. Chapter 11 deals with the consequences.
The six things to remember
| What the system does | What it means for you |
|---|---|
| Decides whether to search at all | Most questions never reach your website |
| Splits one question into many | You compete on questions nobody typed |
| Retrieves passages, not pages | Sections have to stand alone |
| Compares passages head-to-head | Being different beats being polished |
| Counts how many sources agree | Third-party corroboration is arithmetic |
| Generates probabilistically | One measurement is not a measurement |
Chapter 4 — The Room You Have to Be In

There is a comfortable story about AI search that goes like this: the old rules are dead, domain authority no longer matters, and a small specialist with well-structured content can leapfrog an incumbent by writing the right passages.
Parts of that are true. The whole of it is not, and the difference is worth a great deal of money.
What the strongest evidence says
Start with the Washington University audit of real AI Overviews. Of the domains cited in those answers, 29.8 percent did not appear anywhere on the accompanying first page of results.
Which means roughly seventy percent did.
Kevin Indig’s analysis of about 1.2 million ChatGPT responses found the same relationship from another angle: pages ranking first in Google had a 43 percent citation rate, three and a half times the rate of pages ranking beyond position twenty.
And when researchers built a benchmark specifically to test whether AI-era content tactics work — presented at NeurIPS in 2025 — they found that most of them were ineffective or actively harmful to ranking, and that traditional search optimization was significantly more effective.
The largest controlled experiment in the field, 252,000 trials across six models, tested eighteen content factors one at a time. The factors that dominated were topic relevance, the presence of price information, recent timestamps, and position in the candidate list. Formatting and content structure showed minimal effect.
The uncomfortable conclusion
Getting into the retrieval set dominates everything that happens afterward.
If you are not findable, not indexed, not ranking, not known — no amount of passage engineering rescues you. The clever stuff operates on the candidates that made it into the room. It does not get you into the room.
This is less exciting than most of what is sold under this heading. It is what the data supports.
It also has an immediate budget implication. An AI visibility program that neglects the fundamentals of being findable is optimizing the second half of a race it hasn’t entered. If your organic search foundation is weak, fixing that is your AI visibility strategy for the next two quarters, and anyone selling you something else is selling you something else.
What genuinely is new
So is any of this new? Yes. Four things, and they’re the four this book spends its middle on.
The entity layer. A machine cannot apply a credibility judgment to something it can’t identify. And because most questions never trigger a search, what the model already believes about your company is the whole game for the majority of queries. That’s a brand and communications problem, not a publishing one, and almost nobody owns it.
The corroboration layer. Grounding confidence counts how many independent documents verify a claim. Independent agreement isn’t reputation management in the soft sense — it’s an input to a score. The correlational data agrees: unlinked brand mentions track AI visibility far more strongly than backlinks do.
The specificity layer. Query fan-out means the competition happens on attribute intersections that never appear in a keyword tool. Winning “best payroll software” is worth less than winning “payroll for a dental practice with tipped employees.”
The agent layer. A brand an assistant can’t read, price, or transact with is a brand it routes around. This is mostly unglamorous data work, and it’s specified in public documentation, which makes it the most certain investment in the book.
The five levers
Which gives us the structure for what follows. Five levers, in dependency order, because they build on each other.
| Lever | The question | Who owns it |
|---|---|---|
| Entity | Can the machine name you? | Brand and communications |
| Reachability | Can it read what you published? | Engineering |
| Corroboration | Does anyone independent agree? | PR and community |
| Specificity | Do you answer the exact question? | Content, briefed by sales |
| Usability | Can an agent act on your behalf? | Ecommerce and product |
Notice the right-hand column. Only two of the five sit where AI visibility usually gets assigned, which is a large part of why it usually stalls.
Entity comes first because nothing else works if the system can’t tell who you are. Usability comes last because it only matters once you’re being recommended.
Let’s start at the beginning.
PART TWO — THE FIVE THINGS THAT DECIDE IT
Chapter 5 — Be Someone It Can Name
A mid-market software company I’ll leave unnamed spent eighteen months repositioning. New category, new messaging, new website. Beautiful work.
Eight months after launch, they asked an AI assistant what their company did. It described the business they had been three years earlier.
Not because the model was stale. Because the model was right. Their LinkedIn page still carried the old description. So did Crunchbase. So did four industry directories, two conference speaker bios, and the boilerplate at the bottom of every press release they’d issued since 2019 — which their PR agency had copied forward, faithfully, for six years.
The website said one thing. Eleven other places said another. The system did what it is built to do: it went with the consensus.
Why this is lever one
Chapter 3 established the fact that reorganizes everything: most questions never trigger a search. For those, the answer comes from what the model already holds — which is a function of how consistently and how widely your company has been described across the web, over years.
The correlational evidence points the same direction. Ahrefs studied roughly 75,000 brands and measured how strongly various signals track AI visibility:
| Signal | Correlation with AI visibility |
|---|---|
| YouTube mentions | 0.71 – 0.74 |
| Unlinked brand mentions across the web | 0.66 – 0.71 |
| Branded anchor text | 0.51 – 0.63 |
| Branded search volume | 0.35 – 0.47 |
| Domain Rating | 0.27 – 0.33 |
| Backlinks | 0.22 – 0.40 |
| Number of pages published | 0.17 – 0.19 |
Two things jump out. Being talked about beats being linked to. And publishing more is almost irrelevant — the weakest signal on the list is content volume, which is where most budgets go.
The distribution is brutal, too. Brands in the top quartile for web mentions averaged 169 AI Overview mentions. The next quartile down averaged fourteen. The bottom half registered between zero and three.
One honest caveat, which the researchers state themselves: correlation is not causation, and there’s an obvious confound. Big companies have more mentions, more YouTube presence, more branded search, and more AI visibility — all downstream of being big. Nobody has run the controlled experiment. What the numbers establish is the shape of the thing, not the mechanism.
The consistency audit
Find out what the ecosystem currently believes about you. This takes about a week and it’s the highest-yield week in the whole program, because it nearly always surfaces contradictions nobody knew existed.
First, write down the canonical facts. One page, signed off by whoever owns positioning:
Legal name, trading name, every former name
A one-sentence category statement: X is a [category] that [does what] for [whom]
Founded date, headquarters, employee band, ownership status
Product names and their categories
Founder and executive names with titles
Official domain, and any other domains you own
Then check every place your company is described against it. Homepage and About page. Your Organization schema. LinkedIn. Crunchbase. Wikidata. Wikipedia, if you have an article. Google Business Profile. Industry directories — G2, Capterra, Clutch, trade associations. Review platforms. App store listings. Your executives’ own LinkedIn profiles. And the press boilerplate on your last ten releases.
That last one catches more errors than any other check on the list.
Score it by counting contradictions, not properties. A category described four different ways across nine sites is four contradictions. That’s a project with an end, which makes it fundable.
Building the reference point
Somewhere has to be the definitive statement of what your company is. Usually the homepage or the About page. It needs four things.
Plain text identity. Somewhere in the HTML, in a sentence a machine can lift: “Acme Logistics is an enterprise supply chain software company serving mid-market manufacturers in North America.” Not a tagline. Not a video. Words.
Brands resist this because it reads flat next to the copy the site was designed around. Put it in the first paragraph of the About page if the homepage can’t carry it. But it has to exist, in text, in the initial HTML.
Organization schema with a complete sameAs. This is the one piece of structured data worth building regardless of anything else in this book, because its job isn’t citation — it’s disambiguation. The sameAs array is your machine-readable assertion that all these scattered profiles are the same company, which is exactly the puzzle the knowledge graph is trying to solve on its own.
{ "@context": "https://schema.org", "@type": "Organization", "name": "Acme Logistics", "legalName": "Acme Logistics Holdings, Inc.", "url": "https://example.com/", "description": "Enterprise supply chain software for mid-market manufacturers.", "foundingDate": "2016-04-12", "sameAs": [ "https://www.wikidata.org/wiki/Q00000000", "https://www.linkedin.com/company/acme-logistics", "https://www.crunchbase.com/organization/acme-logistics", "https://www.youtube.com/@acmelogistics" ] }
Maintenance. A stale reference page teaches the system your facts are unreliable. Quarterly review, named owner.
Identities for your people. If your expertise argument rests on named humans, those humans need resolvable identities: consistent bylines, author pages with real credentials, author markup pointing at them, matching LinkedIn and speaker profiles. An expert the system can’t identify contributes nothing.
Wikidata: an hour well spent
Wikidata’s bar is much lower than Wikipedia’s. Wikipedia wants significant coverage in multiple independent secondary sources. Wikidata wants a clearly identifiable entity describable with serious public references. Most real businesses qualify for Wikidata. Most do not qualify for Wikipedia.
The process takes about an hour. Search first, using your exact legal name, because duplicate items are common and annoying to merge. Create an account under a real name or a clearly branded handle, and disclose any paid relationship. Add your label, a short disambiguating description, and aliases including former names. Save it and record your Q-identifier — that’s the durable machine-readable handle for your company, and it belongs in your sameAs.
Then add referenced statements: what kind of organization, country, headquarters, inception date, founder, official website, industry, and external identifiers.
The failure mode is simple. Statements without references get reverted. Your own blog and your own press releases don’t count as references. Thin entries get deleted, and a deleted item is harder to rebuild than a good one is to create.
Wikipedia: read this before your agency pitches you
Most companies don’t qualify, and pursuing an article anyway is an active risk.
Paid advocacy is forbidden. Paid editing must be disclosed — employer, client, affiliation — and failing to disclose violates the Wikimedia Terms of Use. You are not supposed to edit the article directly; the sanctioned route is a request on the talk page with full disclosure, which may simply be declined. Getting it wrong can mean account blocks, exposure under FTC guidelines and European fair-trading law, a press cycle about the attempt, and permanent loss of control over what the article says.
Notability for companies requires significant coverage in reliable secondary sources independent of you. Funding announcements, press releases, routine trade coverage, and interviews with your own executives generally don’t count.
An agency promising to get you a Wikipedia page is selling you either a terms-of-use violation or a deletion debate. Fund Wikidata, which you can legitimately build, and let Wikipedia follow real notability if it ever arrives.
There’s an irony here worth noticing. Wikipedia is the most-cited domain in AI answers, somewhere between 5 and 13 percent of ChatGPT citations depending on the study. Its own human pageviews fell about 8 percent year over year, which the Wikimedia Foundation attributes partly to generative AI. The most valuable source in the answer economy is being drained by it.
The entity checklist
Write the canonical facts page and get positioning sign-off — 1 day
Audit every external property against it; count contradictions — 1 week
Fix them, starting with LinkedIn, Crunchbase and press boilerplate — 2–4 weeks
Put a plain-text identity statement on the homepage or About page — 1 day
Ship Organization schema with a complete sameAs — 1 day
Create a referenced Wikidata item; record the Q-ID — 1 day
Build author pages for your named experts — 1 week
Brief the PR agency; update the boilerplate everywhere — 1 day
Put a quarterly consistency review on someone’s calendar — ongoing
None of this is expensive. Most of it has been sitting undone in every organization I’ve looked at, because it belongs to nobody in particular.
Assign it.
Chapter 6 — Be Reachable
This is the least glamorous lever, the one with the strongest evidence behind it, and the one where a single wrong line in a configuration file can remove you from a platform entirely without anyone noticing for months.
The bots that matter, and the one everyone confuses
The most commonly botched item in this whole field is the difference between the crawlers that gate training and the crawlers that gate citation. Block the wrong one and you either hand over training data you meant to withhold, or you delete yourself from a platform’s answers.
| Platform | Bot | What it does | What blocking it costs you |
|---|---|---|---|
| OpenAI | GPTBot | Training | Excludes you from training. Does not affect ChatGPT search citation |
| OpenAI | OAI-SearchBot | Powers ChatGPT search | Removes you from ChatGPT search answers |
| Anthropic | ClaudeBot | Training | Excludes future content from training data |
| Anthropic | Claude-SearchBot | Improves Claude search | Reduces visibility in Claude’s search results |
| Perplexity | PerplexityBot | Surfaces and links sites; not used for training | Removes you from Perplexity |
| Googlebot | Search, News — and AI Overviews and AI Mode | Removes you from Google entirely | |
| Google-Extended | Not a crawler. A permission token for Gemini apps | Does not affect AI Overviews or AI Mode |
That last row deserves a paragraph of its own, because the misconception is everywhere.
There is no bot called Google-Extended fetching your pages. Googlebot crawls as it always has. The token tells Google whether the resulting content may be used to train and ground its Gemini products. Google’s documentation says plainly that it does not affect inclusion in Google Search and is not a ranking signal. AI Overviews and AI Mode are features inside Google Search. So blocking Google-Extended does not touch them. Teams have spent quarters believing otherwise.
And a trap worth stating plainly, because it’s the most likely way a reader breaks their own site while trying to follow this chapter: robots.txt allows by default. You do not need to “explicitly allow” anything. Worse, creating a named group — User-agent: OAI-SearchBot — means that crawler reads only that group and ignores your User-agent: * rules entirely, silently unblocking everything you had disallowed globally.
The correct action is to verify these bots aren’t disallowed. Not to add groups for them.
The deadline on your calendar
If you sit behind Cloudflare, this section has a date in it.
In July 2026 Cloudflare replaced its single AI-bot toggle with three behavioral categories: Search (indexing your content to answer questions about it later), Agent (acting in real time for a person), and Training. Allow, block, or block only on pages with ads — per category, on every plan including free.
From 15 September 2026, for new domains, new sites on existing accounts, and free-tier customers: on pages that display ads, Training and Agent are blocked by default. Search stays allowed.
The trap is the interaction. Crawlers that combine search and training get blocked if training is disabled, because the most restrictive rule wins. A site that monetizes with advertising and accepts the new defaults can quietly lose AI-search visibility from mixed-purpose crawlers, with no error message and no obvious symptom.
Audit your zone settings.
The economics driving all this are worth a line, because they explain why access is getting harder. By mid-2026, 57 percent of web traffic was bots — the first time automated requests exceeded human ones. The crawl-to-referral ratios are extraordinary: some AI crawlers fetch thousands of pages for every visitor they send back, against roughly five to one for Google. Publishers noticed. The rules are being renegotiated in public.
Render it on the server
Here is the least contested finding in this entire field, measured independently, with no dissent:
Most dedicated AI crawlers do not execute JavaScript.
Vercel, measuring around 1.3 billion monthly AI crawler requests, found no major AI crawler running JavaScript. They fetch JS files. They never execute them. A separate analysis of 23 crawlers found 69 percent unable to execute JavaScript, with Googlebot, Bingbot and Gemini’s live fetch the exceptions.
So: anything you want cited has to be in the initial HTML response. Not after hydration. Not in a tab that loads on click. Not behind a lazy-load. Not in a JavaScript-injected accordion.
The first test is view-source, not the DevTools inspector — the inspector shows you the rendered DOM after JavaScript has run, which is exactly what these crawlers never see. The definitive test is curl with the bot’s own user-agent string, since edge logic and user-agent-based rendering can serve a crawler something quite different from what your browser gets.
Both measurements are a year or more old now, in the fastest-moving corner of this subject. Re-test rather than assuming.
Stop wasting the crawl
The same Vercel data found something cheap to fix. ChatGPT’s crawler spends nearly 35 percent of its fetches on 404s. Claude’s spends 34 percent. Googlebot spends 8 percent.
Roughly a third of the AI crawl budget spent on the average site is burning on dead URLs.
Pull your server logs, filter by AI user agent, sort by status code, fix the worst offenders. It’s an afternoon, and it directly increases the share of your real content that gets retrieved.
The new traffic class nobody has a policy for
A different kind of visitor has appeared, and most security teams have never discussed it.
Browser-based agents — Perplexity’s Comet, OpenAI’s Atlas, Claude’s Chrome extension — made up roughly 71 percent of observed agentic activity in a 2026 measurement. Nearly 70 percent of that activity touched product and search pages. Only 3 percent touched checkout.
The practical risk: these look like browsers, not declared crawlers. Aggressive bot mitigation blocks them. When it does, you aren’t blocking a scraper — you’re blocking a customer’s assistant halfway through a task.
Get your security and marketing teams in a room about this once, deliberately, before it happens rather than after.
On structure, and what the evidence actually supports
Now the part where I have to be more careful than most writing on this subject, because the evidence cuts against the received wisdom.
Retrieval works on chunks. That’s documented. But the largest controlled experiment found formatting and content structure had minimal effect across every model tested, and the peer-reviewed benchmark found conversational rewrites frequently hurt ranking.
So structure your content well because it makes retrieval mechanically plausible and because it’s good writing. Not because there’s evidence of a large visibility lift. Specifically:
Don’t orphan your sections. A block that opens “as mentioned above” is ambiguous alone. Restate the subject in the first sentence. Cheap, sound.
Front-load the answer. Citations peak in the first fifth of a page and fall to almost nothing in the last tenth.
Organize around real questions. The one well-evidenced structural finding is that question-shaped queries trigger AI Overviews nearly seven times more often. That’s about how people ask, not how you format — but building content around actual questions is how you meet those queries.
Don’t fund a reformatting project. If someone proposes restructuring four hundred pages into bullet lists for AI visibility, ask for the evidence.
One button not to press
Google added a Search Console toggle that removes your site from AI Overviews, AI Mode and Discover AI features while keeping you in classic search.
Three reasons to leave it alone. You can’t make the decision on data, since Search Console reports AI impressions but not clicks or CTR — you’d be trading an unmeasured benefit for an unmeasured cost. It may take more than you intend, since a meaningful share of trending news queries embed Top Stories carousels inside AI Overviews. And in a comparative system, removing yourself from the candidate set doesn’t reduce your competitor’s visibility. It increases it.
The reachability checklist
Check robots.txt against the bot table — verify, don’t add groups — 1 day
Audit Cloudflare and WAF settings before 15 September 2026 — 1 day
Confirm key content is in view-source, then confirm with curl and each bot’s user agent — 2 days
Move client-side-only content to server-side rendering — varies
Pull AI-bot logs; fix the worst 404s and redirect chains — 1 day
Confirm no noindex or nosnippet on pages you want in answers — 1 day
Review bot mitigation against agentic browsers — 1 day
Don’t press the opt-out button
Chapter 7 — Be Backed Up
Chapter 3 established the mechanism: how many independent documents verify a claim is an input to whether it survives. Chapter 5 established the correlation: being talked about tracks AI visibility far more strongly than being linked to.
These are two separate arguments, and neither proves the other. Both point at the same work — which is the most expensive lever, the slowest, the least controllable, and the one with the highest ceiling.
Where answers actually come from
The distribution of sources is not what most people assume, and it varies enormously by platform.
Wikipedia is the most-cited single domain in ChatGPT, taking somewhere between 5 and 13 percent of citations depending on the study and period. Reddit is first or second nearly everywhere: Ahrefs found it the single most-cited domain in Gemini at 29 percent of top-fifty mention share, with YouTube second. In Perplexity, YouTube leads at 31 percent, Reddit second.
But concentration is lower than the headlines suggest. Profound’s analysis of roughly 730,000 real ChatGPT conversations found the top ten domains accounting for only 12 percent of all citations. The tail is long.
Now the finding that should change how you allocate. BrightEdge compared five engines across nine industries and measured overlap two ways:
Which sources they cite: 16 to 59 percent overlap.
Which brands they recommend: 36 to 55 percent overlap.
The engines disagree substantially about where to pull information from. They agree much more consistently about which brands belong in the answer.
Kevin Indig found the same thing from another direction: 91 percent of citations appear on only one of ChatGPT, Perplexity, or AI Overviews.
The strategic instruction falls straight out of that. Chasing per-platform citation tactics is a treadmill. Brand-level outcomes converge. Optimize the entity and the corroboration, not the platform.
Being cited is not being recommended
Bring this one to the first meeting where somebody shows you an AI visibility dashboard.
Semrush and Kevin Indig examined 3,981 domain appearances across 115 prompts, fourteen countries and four engines. Sixty-two percent were “ghost citations” — the site linked as a source, the brand never named in the answer. Only 13 percent got both.
The platform split is nearly inverse. ChatGPT cites your domain 87 percent of the time and names your brand 21 percent. Gemini cites 21 percent and names 84 percent.
These are two different outcomes and they need two different metrics. Being the footnote is not being the recommendation. If your dashboard reports one number called “visibility,” find out which one it’s measuring.
What earns citations
The largest study available on this comes from Muck Rack, which analyzed over a million links from AI responses across ChatGPT, Claude and Gemini over six months. They sell PR software, so read it as interested research — but it’s the best-documented dataset in the category.
Earned media accounted for 82 percent of all citations (revised upward to 84 percent in a later update). Journalistic sources took roughly a quarter of all links. Half of all citations went to content published within the previous eleven months, and about 4 percent to content published in the previous week.
Seer Interactive triangulates the recency finding from server logs: around 65 percent of AI bot hits target content published in the past year, 89 percent within three years, only 6 percent older than six.
Recency is a retrieval factor, and it’s one of the few on this list you control directly. A refresh program that genuinely updates and re-dates evergreen content is doing retrieval work, not housekeeping.
Where commercial queries are decided
For “best X” and “X versus Y” — the queries nearest to revenue — the evidence is consistent and unwelcome.
Listicles are the largest single page type by citation share, around 22 percent. Of those, roughly 80 percent are third-party listicles. Only 20 percent are brand-authored. Comparison pages, counterintuitively, take under 3 percent; the engines appear to prefer pre-consolidated recommendations to side-by-side tables.
The concentration is severe: about thirty domains capture two-thirds of citations within a topic, and in product-comparison topics the top ten capture nearly half.
In B2B software specifically, an analysis of 57 million citations across fifty companies found Reddit at 28 to 31 percent, YouTube at 15 to 20, LinkedIn at 8 to 15 — and G2 at 11 percent on branded queries but out of the top fifteen entirely on unbranded ones. Brand-owned content was a small fraction of everything cited.
That G2 result is worth sitting with. Review sites show up when someone names your brand and vanish when they don’t. That’s a reputation-defense job, not a discovery job, and they deserve different budgets.
The practical implication: for commercial queries, the highest-leverage work isn’t on your website at all. It’s being accurately represented in the third-party roundups and category comparisons the engines actually cite. That is a PR and analyst-relations program, not a content program.
The most underrated surface here
Two independent measurements point at YouTube, and almost nobody’s strategy reflects it.
BrightEdge found YouTube averaging around 20 percent citation share across AI platforms — and 29.5 percent inside Google AI Overviews, where it is the single most-cited domain. Vimeo and TikTok sit at 0.1 percent each. This isn’t “video.” It’s YouTube.
Ahrefs found YouTube mentions the strongest single correlate of AI visibility across every platform they tested, higher than brand mentions, far higher than backlinks.
The two datasets disagree sharply on Perplexity — BrightEdge puts YouTube at under 10 percent there, Ahrefs at 31 — which is a reminder that these are different measurements with different denominators, not one finding confirmed twice. Causality is unestablished and the confound is obvious.
Still: the platform is owned by the company running the largest answer engine, the correlation is the strongest in the dataset, and the cost of testing is low. Fund an experiment.
The line you must not cross
Community platforms matter, which makes the temptation to manufacture presence on them large. Don’t.
The legal position is no longer ambiguous. The FTC’s rule on consumer reviews and testimonials took effect in October 2024 and prohibits, among other things: fake reviews including AI-generated ones; reviews compensated on condition of sentiment; undisclosed reviews by employees or their immediate relatives; company-controlled sites posing as independent; suppressing negative reviews through threats; and buying fake engagement.
Civil penalties reach $53,088 per knowing violation, and “per violation” can mean per review. Enforcement is live: a 2026 settlement over employees posing as users and incentivized five-star reviews produced a $4 million judgment.
Where the line actually sits is more nuanced than the panic suggests. Paying for reviews isn’t automatically illegal — conditioning payment on sentiment is. “Tell us how much you loved your visit and get a $5 coupon” violates the rule. “Tell us about your visit and get a $5 coupon” doesn’t. General solicitations to customers are safe even if employees respond, as long as no sentiment requirement is attached. Employees may review if they clearly disclose the relationship.
The reputational position is worse than the legal one. A marketing agency publicly boasted in 2025 about running forty-plus fake Reddit accounts staged as authentic player discoveries; the post was deleted inside a day and both the agency and its client issued apologies. When researchers ran undisclosed AI-generated personas on Reddit for months, Reddit’s chief legal officer called it “deeply wrong on both a moral and legal level,” banned the accounts and issued formal legal demands to the university involved. The paper was withdrawn.
If Reddit will pursue a university, it will pursue you.
The asymmetry is the whole argument. The upside of manufactured community presence is a temporary bump in a metric nobody can measure reliably. The downside is a statutory penalty per item, a permanent negative record that AI systems will retrieve for years, and a press cycle. There is no version of this trade that works.
What legitimate participation looks like
In descending order of defensibility:
Answer questions in your own name, with your affiliation visible. Reddit’s policy targets repeated unsolicited promotion, not participation. A named employee giving a genuinely useful answer is within the rules and within the norms.
Host AMAs through official channels, coordinated with moderators.
Publish on YouTube. Given everything above, this is the highest-leverage community surface available and it carries no astroturfing risk at all.
Solicit reviews generally, without sentiment conditions, from actual customers.
Fix the substance. Product and service problems become AI-visible problems, because the community discussion becomes the retrieval corpus. No communications strategy outruns this.
The consensus test
The deepest point in this chapter is structural.
A system seeking corroboration finds three narratives about you. Owned — your website, blog, releases. One voice. Earned — press, analysts, third-party roundups. A second. Community — Reddit, reviews, forums, YouTube. A third.
If owned and earned claim excellence while the community reflects frustration, the system synthesizes toward the community signal. Not because it has taste, but because its architecture favors multi-source agreement over single-source claims.
The brands that win are the ones where the story on the website, the story told by third parties, and the story told by actual customers all point the same way.
That is not a marketing problem. It is an operating one, which is why this lever belongs to the executive team rather than the content calendar.
The corroboration checklist
Map the third-party roundups and review pages cited for your top ten commercial queries — 1 week
Correct factual errors about you in them — the fastest win available — 2–4 weeks
Build analyst and journalist relationships with the outlets that actually get cited, not the prestigious ones — ongoing
Publish original research nobody else has; make the data quotable — quarterly
Fund a YouTube program with real substance — ongoing
Put content on a refresh cycle; recency is retrieval — ongoing
Establish a named, disclosed presence on the two forums where your category is discussed — ongoing
Run a sentiment-neutral review solicitation program — ongoing
Put the FTC review rule in every agency contract — 1 day
Never buy, seed, or sentiment-incentivize anything
Chapter 8 — Be Specific
Traditional search rewarded going broad first. Build category pages for high-volume terms, accumulate authority, let it lift everything beneath.
Answer engines invert this, for a mechanical reason rather than a philosophical one.
When someone asks a specific question, the system breaks it into sub-questions targeting each attribute. A company with content addressing exactly that intersection wins retrieval for it, largely regardless of overall authority. In the vector space where this happens, you compete on proximity to a specific question, not on aggregate reputation.
That’s the opening for specialists, and it’s real. A niche company with authoritative, data-rich content on a narrow domain sits closer to a specific question than a generalist whose coverage is broad and thin. Large companies’ legacy content libraries — optimized for high-volume head terms — often produce few passages that survive chunking and comparison against a precise attribute combination.
But hold that against Chapter 4. Roughly 70 percent of AI Overview citations come from domains already ranking on page one. Specificity is the differentiator within the retrievable set. It is not a substitute for being in it.
The honest formulation: authority gets you into the room. Specificity wins the argument inside it. You need both, and most organizations have exactly one.
Building a question inventory
This is the core new workflow. Two weeks the first time, a day per quarter after.
Step one: harvest real language. Not invented language. In rough order of value:
Sales call recordings — the questions prospects actually ask, in their own words. This is the richest source in most organizations and nobody in marketing has ever listened to it. Then support tickets and chat logs. Your site’s internal search. Sales objection logs, which are comparison questions in disguise. Community threads in your category. Search Console queries, still useful, now one input among several.
Step two: build the attribute matrix. For each core buying question, list the constraints a real buyer applies — company size, industry, integrations, budget, deployment model, compliance regime, sophistication, the specific job to be done. Then generate the intersections.
Not “best project management software,” but:
project management software for a 15-person construction firm that needs offline access
project management tool that works with QuickBooks and doesn’t charge per seat
simplest project management software for a team that has never used one
You aren’t writing a page for each. You’re mapping the sub-question space you’re being evaluated against.
Step three: classify.
| Type | Example | Priority |
|---|---|---|
| Category definition | “what is X” | Low — you’ll lose to Wikipedia |
| Comparison | “X vs Y” | High — commercial, and produces more brand mentions |
| Constrained recommendation | “best X for [situation]” | Highest — the sweet spot |
| Implementation | “how do I do X with Y” | High — proves experience |
| Troubleshooting | “why does X do Y” | Medium — strong experience signal |
| Brand-specific | “is [you] good for X” | High — reputation defense |
Step four: score for winnability. Three checks per high-priority question. Who currently gets cited — run it several times and record the domains. Do you rank on page one for the closest search equivalent; if not, that’s your first job. And: do you have information nobody else has for this question?
That third test kills more content plans than any other, and it should.
Information gain
Chapter 4 was hard on content tactics. Formatting effects measured near zero. Rewrites sometimes hurt. But one thing survived in every study that found anything at all: content carrying information a competing passage cannot.
In the academic work on this, the three highest-performing interventions were adding quotations from recognized authorities (+41 percent), adding statistics (+31 percent), and citing sources (+27 percent). Keyword stuffing went backwards, at −8 percent — the worst-performing method tested. In the largest controlled experiment, the presence of real price information and evidence-backed claims carried some of the largest effects measured.
The mechanism is the head-to-head comparison from Chapter 3. Faced with two passages, the system prefers the one offering something the other doesn’t. A verifiable number. A named source. A real price. A dated measurement. A documented case. A passage that restates common knowledge in slightly different words offers nothing to prefer.
In practice:
Publish real numbers. Your data, your benchmarks, your pricing. If you won’t publish price, you’re absent from one of the strongest content factors anyone has measured.
Name your sources. Attribution is a retrieval asset, not an academic courtesy.
Date everything. An undated page is an old page.
Publish specifications. Missing specs measured as a significant negative. Tables of real attributes beat prose about benefits.
Write from experience. Case studies, first-hand accounts, practitioner data. As the web fills with generated content, evidence of actual involvement appreciates.
Be direct. Hedged writing measured worse than confident writing. That’s not license to overclaim. It’s license to state what you know plainly.
One caution: these advantages compress as competitors adopt them. Information gain isn’t a permanent moat. It’s a moat exactly as long as you actually have information others don’t — which is an argument for original research over content production.
The build order
Tier one: constrained recommendation pages. The exact intersections your best customers describe. Real specs, real prices, real constraints, and honest statements about who this isn’t for. Unglamorous, invisible in a keyword tool, and they win sub-questions.
Tier two: comparison content, including against competitors, written fairly enough that the system will use it. Fairness is instrumentally optimal here — a comparison that reads as marketing gets discounted during grounding, because it contradicts everything else the system found.
Tier three: implementation and troubleshooting depth. The material that proves experience, and the material people link to in forums.
Tier four: original research. The most durable asset, and the one most likely to be cited by the third-party sources from Chapter 7. One well-designed annual study beats fifty blog posts.
Tier five: category education. Last. You’ll lose “what is X” to Wikipedia, and it’s the traffic least likely to convert.
Three things not to do
Don’t mass-produce permutation pages. Generated combinations fail on information gain by construction, and content volume is the weakest correlate on the entire list. Fifty thin pages lose to five substantial ones and cost you crawl budget you’re already wasting.
Don’t restructure your library for “chunkability.” The evidence isn’t there.
Don’t write for the machine at the reader’s expense. The benchmark finding that conversational rewrites often hurt ranking is the empirical version of something that should already be obvious. These systems are trained on human preference. Writing that reads as machine-directed is a detectable signal, and detectable signals get discounted.
Chapter 9 — Be Usable
There is a lot of noise about agentic commerce and considerably less production reality. Both matter, and the gap between them is where budgets get wasted.
What’s actually happening
The forecasts are large. McKinsey projects that by 2030 the US retail market alone could see $900 billion to $1 trillion in “orchestrated revenue” from agentic commerce. Morgan Stanley projects $190 to $385 billion in US e-commerce spending from agentic shoppers by the same year.
Two things before either number goes on a slide. McKinsey’s “orchestrated revenue” counts any purchase journey an agent coordinates or influences — far broader than agent-completed checkouts. Morgan Stanley’s $385 billion is the top of a range and covers US e-commerce only. Placed side by side they look like independent confirmation of each other. They aren’t; they differ threefold mostly because they measure different things.
Meanwhile, in the actual market: in March 2026 OpenAI abandoned native in-ChatGPT checkout and refocused on product search and discovery, with purchases completing in retailer apps. The stated reason was that people research products in ChatGPT but don’t buy there, and merchant adoption was minimal — Shopify’s president said roughly a dozen Shopify merchants were using the AI checkout tools despite integrations being available across three platforms.
Independent measurement agrees. Of observed agentic browser activity, nearly 70 percent touched product and search pages. Just over 3 percent touched checkout.
AI commerce today influences discovery, not conversion. Budget accordingly. An “agentic commerce transformation” is funding a future. A clean product feed is funding the present.
The protocol map, briefly
Four acronyms circulate. Here’s what each is and whether it belongs on your roadmap.
| Protocol | Who | What it does | On your roadmap? |
|---|---|---|---|
| UCP — Universal Commerce Protocol | Agent-to-retailer commerce, end to end | Yes, if you sell online | |
| ACP — Agentic Commerce Protocol | OpenAI and Stripe | Buyer–agent–merchant checkout | Watch; sequence behind UCP |
| MCP — Model Context Protocol | Anthropic-originated, open | Connects models to tools and data | Indirectly — it’s the plumbing |
| A2A — Agent2Agent | Linux Foundation | Agent-to-agent interoperability | No — enterprise interop, not marketing |
UCP is the live one. Google announced it in January 2026, co-developed with Shopify, Etsy, Wayfair, Target and Walmart. By summer, Google Shopping’s Universal Cart had launched in the US across Search and the Gemini app, with UCP-powered checkout live at Nike, Sephora, Target, Ulta, Walmart, Wayfair and select Shopify merchants.
The single most concrete, checkable artifact in this entire space is UCP’s JSON manifest at /.well-known/ucp, which advertises your supported services and endpoints so agents can find your capabilities without a hard-coded integration. Your engineering team can verify in one HTTP request whether you have one. Very few brands do.
The eight attributes worth doing today
If you sell products, this is the most certain AI visibility work available, because it’s specified in public documentation rather than inferred from behavior.
Google connected eight Merchant Center attributes to AI Mode and AI Overviews, rolling out from January 2026:
Product Highlight — short selling benefits, 1–150 characters, two to a hundred per product. Google’s own language: helps customers discover products on AI-driven surfaces.
Product Detail — technical specifications by section, attribute and value.
Variant Option — differences beyond color and size.
Item Group Title — a generic title for a variant group.
Related Product — six relationship types, from “often bought with” to “substitute.”
Question and Answer — up to 1,000 characters each for question and answer, 10,000 combined per product, thirty pairs. Google says it’s “primarily intended for conversational experiences.”
Document Link — up to five PDFs such as manuals, which Google uses to answer detailed questions in AI Mode.
Popularity Rank — a 0–100 sales performance signal.
Note the irony in number six. Google fully deprecated FAQ rich results on websites during 2026. Meanwhile Q&A attributes in your product feed are explicitly built for conversational surfaces. The channel moved, not the content need. If you have a library of good customer questions sitting in deprecated FAQ blocks, they belong in the feed.
Can an agent actually finish a task on your site?
Beyond feeds and manifests, ask this once a quarter. Take an agentic browser, give it a realistic job — find a specific product variant, get a quote, check availability, download a spec sheet — and watch where it fails.
The failure modes are consistent and mundane. Content behind interaction: specifications in a tab that loads on click, pricing behind a configurator, availability behind a form with no URL state. Bot mitigation that challenges the agent and stops it. Unstable URLs, since agents re-fetch and session-dependent carts break on the second visit. Information that exists only inside images. Fulfilment terms written in prose that requires interpretation rather than in structured, machine-readable form.
Every one of those is also a human usability problem, which is the useful part. Agent-readiness work has a business case even if agentic commerce underdelivers.
Fund and defer
Fund now: a clean, complete, regularly refreshed product feed — the common denominator across every platform. The eight attributes, in the order listed. The /.well-known/ucp manifest if you sell where Universal Cart is live. Price, availability and fulfilment terms in the initial HTML. And a quarterly agent-completability test.
Defer: ACP checkout endpoints until OpenAI’s direction settles. And anything sold as an “agentic commerce readiness program” rather than a set of tickets. The work is a feed, a manifest, and a measurement loop. If a proposal is bigger than that, ask what specifically it ships.
PART THREE — RUNNING IT
Chapter 10 — What Actually Works
This field has a research problem. Not a shortage — the volume is impressive — but most of it is produced by companies selling tools in the category, most of it is correlational, and the studies that disagree with each other rarely get read together.
You are going to be asked to approve budget against this. So here is the scoreboard, sorted into what we know, what’s contested, and what’s being sold to you.
Everything below is graded:
| Grade | Meaning |
|---|---|
| A | Multiple independent measurements, or a first-party platform specification. Act on it. |
| B | One good study or consistent practitioner measurement with a plausible mechanism. Act, verify locally. |
| C | Contested. Correlational, confounded, or the studies disagree. Do it if it’s cheap. |
| D | No evidence, or evidence against. Don’t fund it. |
| F | Evidence of harm, or unlawful. Prohibit it. |
A plus or minus modifies within a grade.
What we know
Server-side render anything you want cited. Grade A. Measured independently by two organizations with no dissenting result. Most dedicated AI crawlers cannot execute JavaScript. Content that only exists after hydration doesn’t exist.
Fix your 404s and redirect chains. Grade A. Roughly a third of AI crawler fetches on the average site hit dead URLs. Mechanically obvious, cheap, and directly increases how much of your real content gets retrieved.
Get your crawler configuration right. Grade A. First-party documentation from every platform. The bots that gate training are not the bots that gate citation, and getting it backwards is common and invisible.
Complete your product feed. Grade A. Google and OpenAI both publish specifications. The eight Merchant Center attributes are explicitly connected to AI surfaces by Google’s own documentation. This is the least ambiguous work in the book.
Establish your entity. Grade A−. Low cost, documented mechanism, strong correlational support. Nobody has isolated the effect of fixing a specific inconsistency, which is why it isn’t a straight A.
Earn third-party mentions. Grade B+. The strongest correlate in the largest correlational study, and separately, corroboration count is a documented input to grounding. Two independent reasons pointing the same way, neither of which proves the other.
Publish real prices, specifications and original data. Grade B+. Price presence carried one of the largest effects in the biggest controlled experiment, with the caveat that the experiment ran in a synthetic testbed rather than in Google.
Organize content around real question-shaped queries. Grade B. Question-form queries trigger AI Overviews nearly seven times more often than other query forms. Note carefully: that’s about how users ask, not about how you format your headings.
Refresh content on a cycle. Grade B. Around 65 percent of AI crawler attention goes to content published within the past year. Recency is a retrieval factor.
Invest in YouTube. Grade B−. The strongest single correlate in the brand-visibility data, and the most-cited domain inside Google AI Overviews. Causality unestablished, confound obvious, cost of testing low.
What’s contested
Restructuring content for “chunkability.” Grade C. The largest controlled experiment found formatting and structure effects minimal across every model tested. A peer-reviewed benchmark found conversational rewrites frequently harmed ranking. Write well because it’s good writing. Don’t fund a reformatting project.
Schema markup as a citation lever. Grade C−. This is the most confidently asserted tactic in the field, so the evidence deserves stating plainly.
Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 matched control pages over eight months, using difference-in-differences. The effect on Google AI Overviews was −4.6 percent, small but statistically significant. On AI Mode, +2.4 percent. On ChatGPT, +2.2 percent. Both of the last two are indistinguishable from zero.
A separate mechanical test put eight product prices on a page, distributed across visible HTML, JavaScript-rendered content, JSON-LD only, and hidden markup, then asked each system to find them. No system found prices that existed only in JSON-LD.
And Google’s own documentation says, in as many words, that there’s no special structured data you need to add to appear in these features.
This does not mean abandon schema. It means stop funding it as a citation lever. Structured data still does real work for entity disambiguation, for Bing and Copilot, for classic rich results, and for commerce feeds. “Getting cited in AI answers” isn’t one of its jobs.
“Share of voice” as a stable metric. Grade C. Non-determinism, personalization and arbitrary prompt-set design are unpriced in every vendor’s reporting. Chapter 11 deals with this.
What’s being sold to you
llms.txt. Grade D. Google’s John Mueller, on the record: “no AI system currently uses llms.txt.” A server-log study across roughly 900 domains over seven months recorded 1,227 requests for the file — fewer than seven per day across all sites — of which two-thirds came from a single data broker and a third from Chrome browsers, meaning humans. Zero came from a frontier AI lab’s crawler.
There is one genuine signal: Chrome added an llms.txt audit in 2026, filed deliberately under agentic browsing rather than SEO. That’s where the file may eventually matter — agent tooling, not search visibility.
Publishing one costs an hour and harms nothing. Don’t let it onto a roadmap as a visibility driver, and don’t pay an agency for it.
Blocking Google-Extended to control AI Overviews. Grade D. It does not do that. Google-Extended governs whether content may be used to train and ground Gemini products. AI Overviews and AI Mode are features inside Google Search, which the token does not touch.
Keyword stuffing. Grade D. Measured at −8 percent — the worst-performing tactic in the academic study of content optimization. It doesn’t fail to help. It hurts.
Manipulating recommendations. Grade F. Covered in Chapter 13. It works, briefly, and then it stops working for everyone including you.
The through-line
Read together, the strongest studies say the same thing, and it isn’t what the category sells.
Getting into the retrieval set dominates everything downstream. Seventy percent of AI Overview citations come from domains already on page one. Pages ranking first get cited three and a half times as often as pages beyond position twenty. Traditional search optimization beat purpose-built conversational tactics in the one peer-reviewed head-to-head. Relevance, recency and candidate position dwarfed every cosmetic content factor in the largest controlled experiment.
What’s genuinely additive is the entity work, the corroboration work, the specificity work, and the agent-readiness work. That’s where the incremental return lives, and it’s four things rather than forty.
Chapter 11 — Measuring Without Fooling Yourself
Before any dashboard, tool selection or board slide, somebody in your organization has to say this sentence out loud:
We are being asked to manage a channel whose cost is measurable and whose benefit is not.
Search Console shows AI impressions but not clicks, click-through rate, queries or position. Bing shows citations but not traffic value. Analytics catches some referrals and loses the rest to Direct. And the dominant effect — being recommended in an answer nobody clicks — leaves no trace at all.
Any measurement plan that promises to close that gap is overpromising. What follows closes what can be closed and is honest about the rest, which is the achievable goal.
Rank is gone, and what replaces it
Under classic search there was one results page and one position. Under personalized, non-deterministic generation there is no universal ranking. A brand that appears first for one person may not appear at all for another.
What actually breaks rank as a metric isn’t the averaging. It’s the variance — across people, because of personalization, and across runs of the same question, because of stochasticity. A single observed position tells you almost nothing about either.
So the thing you’re estimating is your average position across the population of people who might ask. You can’t observe that population, so you approximate it with three things: a question set that represents real demand, enough repeated runs to average out randomness, and enough platform coverage to reflect where your buyers actually are.
Miss any of the three and you’re producing noise with invisible error bars.
The sampling requirement nobody meets
Researchers at the University of St. Gallen tested something the vendor category would rather nobody tested: run the same prompt over and over and see how much the answer moves.
Identical prompts on identical days produced source overlap of 0.34 to 0.42 and brand-mention overlap of 0.45 to 0.59. The instability held inside 24-hour windows, so it’s the model, not the news.
Their recommendation: at least seven runs per prompt per day for brand tracking to get the standard error below 0.10, at least eight for source coverage, and rolling windows of two to four weeks. Their conclusion, in their words: single observations are misleading.
Most commercial tools sample far below this, and almost none publish confidence intervals. Three rules follow, and they should be non-negotiable:
Never report a single-day figure. Rolling averages over two to four weeks only.
Ask every vendor in procurement how many runs per prompt per day they execute, and whether they publish confidence intervals. The answer is diagnostic. Most won’t have one.
Treat any before-and-after case study without disclosed repeated sampling as unproven. Including your own. Especially your own.
Two kinds of tool, and why they disagree
Every vendor sells “share of voice in AI answers.” They do not measure the same thing.
Synthetic prompting. The tool fires a prompt list you configure at model APIs or chat interfaces on a schedule, parses the responses, aggregates. Most of the category works this way. You control the question set and get per-prompt diagnostics. The weakness is fundamental: you’re measuring answers to questions you invented.
Real-prompt corpora. A smaller group claims to derive prompts from observed behavior rather than invention — Ahrefs states an index of over 473 million monthly prompts; Profound claims an index built from over 1.5 billion real conversations. Neither publishes an auditable methodology for how they obtain them at that scale, so treat the volume claims as directional rather than audited.
A working approach: buy one of each. They will disagree, and the disagreement is the methodology rather than a bug. Seeing both keeps you honest.
Rough entry pricing, useful mainly for scale: Profound from around $99 a month, Ahrefs Brand Radar from $199, Otterly from $29, Semrush’s AI toolkit around $99 per domain, Scrunch from $250, with Conductor and BrightEdge sold as enterprise. Verify before quoting — these move.
The structural point is that a serious measurement program costs a few thousand dollars a year, not a few hundred thousand. The constraint is analytical discipline, not budget.
The free sources you’re probably ignoring
Google Search Console generative AI reports, launched mid-2026. Impressions from AI Overviews and AI Mode, by page, country, device and date. No clicks, CTR, queries, position or conversions. One trap: impression counts differ by aggregation level, so property-level and page-level totals aren’t comparable.
Bing Webmaster Tools AI Performance report. Total citations, unique URLs cited daily, sample triggering queries, per-page counts. Free, and the only first-party citation count anybody publishes.
GA4’s AI Assistant channel. Sets medium to ai-assistant when it recognizes a referrer. Remember it’s a floor: in-app and mobile traffic often arrives with no referrer at all, and one study found 22 percent of AI Overview traffic misfiled as Direct.
Your server and CDN logs. The most underused source in the stack. Filter by AI user agent and you get crawl frequency by bot, which content is being fetched, how much budget is burning on 404s, and whether your robots.txt and firewall changes actually took effect. Free, and it answers questions no dashboard will.
Designing a question set worth measuring
In synthetic prompting, the question set is the entire ballgame. A tool tracking 25 prompts is a tripwire, not a measurement system.
For a mid-sized program, aim for 120 to 200 questions:
| Segment | Share |
|---|---|
| Constrained recommendation — “best X for [situation]” | 40% |
| Comparison — “X vs Y”, “alternatives to Y” | 20% |
| Brand-specific — “is [you] good for X” | 15% |
| Implementation and troubleshooting | 15% |
| Category education | 10% |
Four rules. Source them from the harvest in Chapter 8, not from a brainstorm. Include questions you currently lose — a set built only from your strengths produces a dashboard that flatters you and teaches you nothing. Include your top three competitors’ strongest positions. And freeze the set for a quarter; when you do change it, version it and treat the series as broken at that point.
What to report
Four numbers, each answering a distinct question. Don’t collapse them into a single “AI visibility score” — those scores hide exactly the trade-offs that matter.
Mention rate. Share of tracked questions where your brand is named in the answer, on a rolling basis. This is the recommendation metric.
Citation rate. Share where your domain appears as a source. Remember: 62 percent of appearances are citations without mentions, and ChatGPT and Gemini are nearly inverse on this. Reporting one and calling it the other is the most common error in the field.
Accuracy. What’s being said, and whether it’s true. This needs a person reading a sample of answers, not a sentiment classifier. It catches problems before they become Chapter 13 problems.
Competitive share. Your mention rate against your named competitors on the same question set. In a comparative system the relative number matters more than the absolute one.
Alongside those, report the inputs you control — entity consistency, share of key content server-rendered, AI crawler 404 rate, third-party roundups where you appear correctly, feed completeness. Inputs move before outputs do, and they’re the only part of the system where cause and effect are legible.
The ROI conversation
You will be asked what this returns. Here’s how to answer without lying.
What you can attribute: AI referral sessions and their conversion behavior, understanding the number is a floor. If that traffic converts materially better than other channels — and the evidence suggests it currently does — that’s a real, defensible, small number.
What you can’t attribute: the influence of appearing in answers nobody clicks. That’s most of the value, and there is no honest way to measure it directly today.
What to do about the gap. Three things.
Proxy it with competitive mention rate. If you’re named in 40 percent of relevant answers and your closest competitor is named in 15, that’s a share-of-consideration position with real economic content even if you can’t price it.
Instrument the qualitative side. Add “did you research this with an AI assistant?” to your demo request and post-purchase surveys. Self-reported attribution is imperfect and it’s currently the only line of sight into zero-click influence. Track it as a percentage over time.
And be explicit about the asymmetry in the business case. The entity and reachability work is small and largely one-time. The corroboration work is real money, but it’s mostly spend you’re already making on PR and content, redirected. You aren’t asking for a new budget line proportional to an unmeasurable return. You’re asking to reallocate an existing one toward mechanisms that are documented.
That last framing is what gets this funded. A request for new money against an unmeasurable return fails. A request to redirect existing money toward better-evidenced mechanisms succeeds.
Chapter 12 — Paying for It
The argument you’ll hear is that AI search requires shifting budget from performance to brand. The evidence for that is better than most claims in this field, and the canonical citation people reach for is weaker than they think. Know both halves.
What the 60/40 rule actually says
The 60:40 brand-to-activation ratio comes from Les Binet and Peter Field’s work for the IPA, drawing on its effectiveness databank. Its authors framed it as a guideline for the average case, not a law.
The critique is legitimate and hasn’t been rebutted on the merits. Byron Sharp of the Ehrenberg-Bass Institute has called the rule “very misleading,” built on “a very weird data set” — the databank is composed of award submissions, which are self-selected, self-reported campaigns that entrants believed had worked. His argument is that the number exists largely because the industry wanted a number.
The honest middle came from James Hurman: the rule may be less scientific than Sharp would like, but few would argue with the observation that marketers tend to underinvest in brand.
And it varies by context — its own authors say so. The optimal brand share runs from roughly 20 to 80 percent depending on how the category buys. For B2B specifically, the analysis of B2B cases in the same databank recommends closer to 50/50, with the authors describing their own findings as tentative and their samples as small.
So the defensible claim is directional, not arithmetic. Most organizations underinvest in brand. The right ratio for your category is not 60/40 because a book said so.
Which way budgets are actually moving
Against you, as it happens, which makes this argument contrarian rather than fashionable.
NIQ’s 2026 CMO survey of over 250 senior decision-makers found only 55 percent allocating 60 percent or more to long-term brand building, down from 59 percent the year before — and only 69 percent saying their C-suite believes in long-term brand value, down from 80 percent. Seventy-four percent report heightened ROI scrutiny.
Gartner’s survey of 401 CMOs found marketing budgets flat at 7.8 percent of revenue, 15 percent of marketing budget allocated to AI generally, 70 percent saying AI leadership is critical, and only 30 percent reporting mature readiness. No discrete line item for any of this.
So you’re making a case for less measurable investment at exactly the moment accountability pressure is tightening. Pretending otherwise won’t help you win the argument.
The argument that actually works
Don’t lead with 60/40. Lead with the mechanism.
One: most questions never trigger a search. Between two-thirds and four-fifths of prompts are answered from what the model already holds. For those, brand prevalence is the visibility mechanism. There’s no content tactic that reaches them. This is a brand argument grounded in architecture rather than in award-show data, and it’s the strongest thing you have.
Two: you can’t buy your way in. Using search as the reference, paid placement and organic recommendation stay structurally separate, because paid bias in organic results destroys the trust that makes the market work — especially in a market where users have demonstrated they’ll switch platforms readily. A platform whose share can fall twenty points in a year can’t afford to sell its recommendations.
State the counterweight rather than hiding it: ads are already appearing inside answers. OpenAI began testing them in early 2026, and by mid-year roughly a quarter of ChatGPT responses contained one, labeled and separated. The organic and paid layers are separating exactly as they did in search. What remains unlikely is paid influence over the organic recommendation.
Three: the correlational evidence points at brand signals, not content signals. Unlinked mentions track AI visibility around 0.66; backlinks around 0.22; content volume around 0.18. Correlation isn’t causation and the confound is real. But if you’re allocating under uncertainty, allocate toward the stronger relationship.
Four, and this is the one that closes. You aren’t asking for new budget. You’re asking to redirect existing spend toward better-evidenced mechanisms. Some content production money moves to original research, which serves both information gain and earned media. Some link-building money moves to digital PR aimed at the outlets that actually get cited, which produces mentions rather than links. Some technical budget moves to server-side rendering and crawl cleanup, which costs less than what it replaces. Entity and feed work is small, largely one-time, and unowned today.
Where the money goes
A defensible allocation for a mid-sized program:
| Area | Share | Why |
|---|---|---|
| Core search fundamentals | ~40% | Getting retrieved dominates everything after |
| Digital PR and earned media | ~25% | Strongest correlate; corroboration is a documented input |
| Measurement and reporting | ~20% | The sampling requirements are real and unmet |
| Enablement and training | ~10% | Most of your team’s mental model is three years old |
| Experimentation | ~5% | YouTube, agent-readiness, whatever comes next |
The line that will get challenged is the 20 percent on measurement. Defend it with Chapter 11: without adequate sampling you can’t tell whether any of the other 80 percent worked, and you’ll spend the difference arguing about noise.
Who owns this
The most common organizational failure here is assigning all of it to whoever owns SEO, when roughly half the work isn’t SEO work.
| Lever | Natural owner | Usual failure |
|---|---|---|
| Entity | Brand and corporate comms | Given to SEO, who can’t change LinkedIn or press boilerplate |
| Reachability | Engineering, with SEO | Given to SEO, who can’t change the rendering architecture |
| Corroboration | PR, analyst relations, community | Given to content, who can’t place earned media |
| Specificity | Content, briefed by sales | Built from keyword tools instead of customer language |
| Usability | Ecommerce and product | Unowned; the feed belongs to whoever built it in 2019 |
What works in most organizations isn’t a new team. It’s a named owner and a quarterly forum: one accountable person, a standing session with engineering, PR, content, ecommerce and legal, and a single dashboard all of them see. The forum exists because the failure mode isn’t incompetence — it’s that five levers sit in five reporting lines and nobody holds the whole picture.
Legal belongs in that room. The review rules, the liability developments, the transparency obligations — all live, and all cheaper to handle as constraints on the plan than as responses to an incident.
Seven questions for an agency
If you’re buying help, these separate the substantive from the performative.
What’s your evidence that this works? In writing.
How many runs per prompt per day does your measurement use, and do you publish confidence intervals?
Are you proposing anything that requires editing Wikipedia? If yes, end the conversation.
Are you proposing anything involving reviews, forum posts or testimonials we wouldn’t want printed in a trade publication? If yes, end the conversation.
What proportion of the work is on our own site? If it’s most of it, they’re selling you Chapter 6 and calling it Chapter 7.
What do you think llms.txt does? Cheap, reliable diagnostic.
What would make you tell us this isn’t working? Anyone who can’t answer has no falsifiable model.
Chapter 13 — When It Goes Wrong
Every chapter so far has been about getting into the answer. This one is about what happens when the answer is wrong, when someone games it, or when a regulator arrives.
The system will state false things about you
This is the likeliest risk to materialize and the least prepared-for.
The Washington University audit verified 98,020 individual claims in AI Overviews against their cited sources and found 11 percent inconsistent — 4 percent actively contradicted by the source, 7 percent simply not present in it. That’s the error rate on claims that carry a citation, which means they look verified.
For your brand this shows up as invented specifications, wrong pricing, wrong availability, misattributed features, discontinued products presented as current, and confident statements about policies you don’t have.
A real case: in 2025 the AI support bot belonging to the developer tool Cursor invented a company policy — that users couldn’t run sessions on multiple machines. No such policy existed. Customers cancelled subscriptions over a rule that was never written, and a co-founder had to state publicly that there was no such policy.
What to do. Monitor for accuracy, not just presence — somebody reads a sample of answers about you every month. Make the correct facts easy to find and hard to contradict, which is the Chapter 5 work with a second justification. Establish a correction path before you need it: who reports an error, to which platform, through which channel, who signs off. And find the upstream source, because most hallucinated brand facts aren’t invented from nothing. They come from a stale directory listing, an outdated release, or a competitor’s comparison page.
Liability is moving, in different directions
In Europe, toward the brand. In May 2026 a Munich court granted a temporary injunction against Google after AI Overviews falsely connected two publishers to scams. The court held Google directly liable, rejected the argument that users should verify independently, and drew an explicit line: an AI overview is Google’s own content, not a list of search results. Since only Google controls the algorithms, Google owns accuracy.
The caveat matters. That’s an interim injunction from a court of first instance, and Google has said it will appeal. It isn’t settled precedent. What it establishes is that the argument works in at least one major jurisdiction, which is a real shift and less than the headlines claimed.
In the US, the other way, so far. A Georgia court granted summary judgment to OpenAI over ChatGPT fabricating embezzlement allegations, on three grounds: a reasonable reader in context couldn’t have understood the output as stating actual facts, no negligence or actual malice, no recoverable harm.
But not settled. A separate case against Google over allegedly fabricated criminal accusations survived a motion to dismiss and moved to discovery.
And you own what your own chatbot says. When Air Canada’s chatbot told a passenger he could apply retroactively for bereavement fares, contradicting the website, the airline argued the chatbot was a separate legal entity responsible for its own statements. The tribunal rejected that outright and held the airline accountable for all information on its site, static or generated.
If you deploy a customer-facing assistant, treat its output as your published statements — because a tribunal already has. Constrain it to verified sources, log what it says, label it.
Manipulation works, and that’s the problem
Researchers at ETH Zürich demonstrated preference manipulation attacks: crafted web content that steers AI-powered search toward the attacker’s product. Manipulated fictional cameras became two and a half times more likely to be recommended, competing successfully against real brands. Fake products moved from a 34 to a 59 percent recommendation rate. Production systems were affected.
Now the finding that should end the discussion internally: when several competitors attack simultaneously, everyone’s recommendation rate degrades. It’s a prisoner’s dilemma. Individually rational, collectively destructive.
Use that argument with anyone in your organization who’s tempted, because it doesn’t require them to share your ethics. The tactic destroys the value of the surface it exploits, and it does so quickly once more than one player adopts it.
The related security exposure is real too, and nobody’s content policy accounts for it. Brave’s researchers demonstrated indirect prompt injection in an agentic browser, and the proof-of-concept vector was a Reddit comment with instructions hidden in a spoiler tag. When a user asked the browser to summarize the page, the AI executed the hidden instructions, accessed the user’s account and exfiltrated credentials.
Your community pages, review sections and user-generated content are now an attack surface that can be turned against your own customers’ assistants. Add prompt-injection review to your moderation policy.
Regulators have arrived
The EU AI Act’s transparency obligations are in force as of August 2026. Users must be told they’re dealing with an AI system unless it’s obvious. Synthetic content, including text, must be marked in machine-readable format — systems already deployed had a grace period into December 2026. Penalties reach €15 million or 3 percent of worldwide turnover.
The direct marketing relevance: AI-generated marketing text distributed in the EU falls within the marking obligation, and AI chat interfaces on brand properties require disclosure. Now, not in some future phase.
The FTC’s “Operation AI Comply” has continued across a change of administration, with more than a dozen cases tied to AI washing in the past year. Two features matter. Enforcement has expanded to B2B marketing — the same substantiation standards apply regardless of audience. And the Commission has held vendors liable for supplying deceptive materials used downstream, which puts your agency and your martech suppliers in scope alongside you.
The review rule from Chapter 7 carries penalties up to $53,088 per knowing violation, with live enforcement since 2026.
The structural risk underneath all of it
Answer engines sit between you and your customers, and their logic is opaque in a way search never quite was. Under classic search, ranking factors were partially understood and results were visible. Now you get impressions without clicks, citations without traffic, and recommendations you can’t observe.
Three specific exposures. Concentration — about thirty domains capture two-thirds of citations within a topic. Volatility — one platform’s share can fall twenty points in a year while another triples. Commercial encroachment — ads inside answers, on a surface smaller than the page it replaced.
The response is deliberate diversification. The entity and corroboration levers are the platform-independent ones; they work across engines precisely because engines disagree about sources and agree about brands. And owned audience relationships — email, community, direct — remain the only fully independent asset.
None of that argues for abandoning AI visibility. It argues against building a business on top of a single opaque intermediary, which is a lesson this industry has already learned once.
PART FOUR — DOING IT
Chapter 14 — The First Ninety Days
Everything in this book, sequenced. Dependencies run downward: entity before corroboration, reachability before content, measurement before any claim of success.
Days 1–30: find out what’s true
The goal of the first month is to know where you stand and to fix what’s silently broken.
| # | Action | Owner | Effort |
|---|---|---|---|
| 1 | Write the canonical facts page; get positioning sign-off | Brand | 1 day |
| 2 | Audit every external property against it; count contradictions | Brand / SEO | 1 week |
| 3 | Check robots.txt against the bot table — verify, don’t add groups | Engineering | 1 day |
| 4 | Audit Cloudflare and WAF settings — deadline 15 September 2026 | Engineering | 1 day |
| 5 | View-source and curl tests: is key content in the initial HTML? | Engineering | 2 days |
| 6 | Pull AI-bot server logs; quantify 404 and redirect waste | Engineering | 2 days |
| 7 | Turn on Search Console generative AI reports and Bing AI Performance | SEO | 1 hour |
| 8 | Verify the GA4 AI Assistant channel; document the Direct gap | Analytics | 1 day |
| 9 | Harvest real customer language from sales calls, support, site search | Content / Sales | 1 week |
| 10 | Baseline accuracy audit: read fifty AI answers about your brand | Marketing | 2 days |
| 11 | Brief legal on the review rule, AI Comply, and EU transparency duties | Legal | 1 day |
| 12 | Name one accountable owner; schedule the quarterly forum | CMO | 1 day |
What you should have at the end: a one-page diagnostic — number of entity contradictions, percentage of key content server-rendered, AI crawler 404 rate, accuracy findings, and whether any platform is currently blocked by misconfiguration.
That last item justifies the whole month more often than not. Blocking OAI-SearchBot while intending to block GPTBot is common, invisible, and entirely self-inflicted.
Days 31–60: fix it and build the baseline
| # | Action | Owner | Effort |
|---|---|---|---|
| 13 | Fix entity contradictions: LinkedIn, Crunchbase, directories, boilerplate | Brand / Comms | 3 weeks |
| 14 | Add a plain-text identity statement to the homepage or About page | Brand | 1 day |
| 15 | Ship Organization schema with a complete sameAs | Engineering | 1 day |
| 16 | Create a referenced Wikidata item; record the Q-ID | SEO | 1 day |
| 17 | Move client-side-only content to server-side rendering | Engineering | 2–4 weeks |
| 18 | Fix the worst AI-crawler 404s and redirect chains | Engineering | 3 days |
| 19 | Build the 120–200 question set from the month-one harvest | Content | 1 week |
| 20 | Buy one synthetic-prompt tool and one real-prompt tool; ask both for run counts | SEO | 2 weeks |
| 21 | Map the third-party roundups cited for your top ten commercial queries | PR / SEO | 1 week |
| 22 | Correct factual errors about you in those sources | PR | 2 weeks |
| 23 | Audit the product feed; populate Product Highlight and Product Detail | Ecommerce | 2 weeks |
| 24 | Review bot mitigation against agentic browser traffic | Security / Eng | 1 day |
What you should have: a measurement baseline with rolling mention rate and citation rate reported separately, and a corrected entity footprint.
Do not report a trend yet. You don’t have one. Reporting a trend from three weeks of data is exactly the error Chapter 11 warns about, and doing it once sets a precedent you’ll spend a year unwinding.
Days 61–90: build the things that compound
| # | Action | Owner | Effort |
|---|---|---|---|
| 25 | Publish constrained-recommendation pages for the top twenty intersections | Content | 6 weeks |
| 26 | Publish prices and specifications wherever commercially possible | Content / Product | 2 weeks |
| 27 | Commission the first original research study | Marketing | starts now |
| 28 | Launch digital PR aimed at the outlets that actually get cited | PR | ongoing |
| 29 | Establish a named, disclosed presence on the two key forums | Community | ongoing |
| 30 | Launch a sentiment-neutral review solicitation program | CX | ongoing |
| 31 | Fund a YouTube pilot with real substance | Content | ongoing |
| 32 | Migrate FAQ content into feed Q&A attributes; attach document links | Ecommerce | 2 weeks |
| 33 | Publish /.well-known/ucp if you sell where Universal Cart is live | Engineering | 1 sprint |
| 34 | Run the first quarterly agent-completability test | Ecommerce | 1 day |
| 35 | Put content on a refresh cycle | Content | ongoing |
| 36 | First cross-functional forum: review inputs and outputs together | CMO | half day |
What you should have: a report with four output metrics on rolling averages, five input metrics you control, an accuracy finding, and an explicit statement of what you still cannot measure.
That last section is the most credible thing in the document. It’s also what protects you when someone asks why the numbers moved and the honest answer is that they didn’t move — they varied.
What not to do in the first ninety days
Don’t restructure the content library for chunkability
Don’t commission a schema program aimed at AI citations
Don’t pay anyone for llms.txt
Don’t block Google-Extended expecting it to affect AI Overviews
Don’t press the Search Console AI opt-out
Don’t seed reviews, forum posts or testimonials — ever
Don’t report a trend before you have properly sampled data
Don’t fund an “agentic commerce transformation.” Fund a feed and a manifest.
If you do nothing else
Check your crawler configuration. Make your identity consistent everywhere it appears. Put your real content in the initial HTML. Start measuring properly.
Four things, all graded A, all cheap, all undone in most organizations right now.
Everything else in this book is optimization on top of them.
Chapter 15 — What to Watch
Forecasting here has a poor record, so what follows is framed as things to watch with thresholds attached, rather than predictions. The useful question isn’t “what will happen.” It’s “what would I do differently, and what would tell me it was happening.”
Will specialists really win?
The theory says yes: because retrieval rewards proximity to a specific question rather than aggregate reputation, small specialists gain an advantage. Small brands double down on niches, large brands respond by segmenting into sub-brands, and a new layer of specialists emerges between manufacturers and consumers.
If that holds, something significant follows: brand equity becomes local to a market rather than a portable asset that lets a big company enter new categories cheaply. That would invert one of the oldest assumptions in marketing strategy.
The counter-evidence is substantial. The largest observational study of brand visibility found a steep incumbency gradient — household names appearing in roughly 73 percent of relevant answers, mid-market brands 44 percent, small and niche brands 11 percent. Whatever theoretical advantage specialists hold, incumbents are winning today.
Watch for: specialists appearing in your constrained-recommendation questions while absent from broad category questions; the incumbency gradient flattening in successive studies; sub-brand launches from large competitors built around narrow use cases.
Either way: the specificity work is the hedge. It costs the same whether you’re the incumbent or the challenger.
Will anyone be able to buy the recommendation?
Probably not the organic one, for the reasons in Chapter 12 — though the separation is thinner than it looks. Ads already appear in about a quarter of ChatGPT responses. Google is piloting programs that let retailers surface exclusive discounts inside AI Mode.
The distinction that matters isn’t “ads or no ads.” That’s settled. It’s whether the organic recommendation can be influenced by spend.
Watch for: unlabeled placement changes correlated with spend; platforms selling “recommendation visibility” as a product; a measurable correlation between ad spend and organic mention rate that survives controlling for brand size. That last one is a study you could actually run.
Do now: measure organic mention rate separately from any paid placement, starting immediately. If the surfaces blur later, you’ll want the clean baseline.
Will agents actually buy things?
The bull case is the trillion-dollar forecasts. The bear case is what happened in March 2026, when OpenAI retreated from native checkout after roughly a dozen merchants adopted it, because people research in the assistant and buy elsewhere.
The likely middle is agent-initiated, human-confirmed: the agent assembles the cart or fills the checkout, and a person confirms. That’s functionally the same as one-click purchase on your own site, it preserves user authority, and it avoids the accidental-purchase problem, where consumers may remain liable for transactions they never explicitly approved. Full autonomy creates regulatory problems nobody has solved.
Watch for: checkout share of agentic traffic rising above single digits; UCP checkout expanding beyond the initial retailer set; a major platform shipping autonomous purchase with a published liability framework.
Do now: the Chapter 9 list and nothing beyond it. A feed, the eight attributes, a manifest, a quarterly test. That work pays off in discovery today regardless of whether transaction volume ever arrives.
Will measurement standardize?
Regulatory pressure says partly. UK competition authorities have required impressions, click-throughs and click-through rate; Google currently supplies only impressions, with further obligations scheduled.
Architecture says no. Personalization makes a universal ranking incoherent. Non-determinism makes single measurement meaningless. No regulator can mandate a number that doesn’t exist.
The likely outcome is partial: better first-party reporting, no standardized share-of-voice metric, continued vendor disagreement. Which means the discipline in Chapter 11 stays a competitive advantage rather than becoming table stakes.
Watch for: Google shipping clicks and CTR in Search Console; an industry body publishing a sampling standard; vendors starting to publish confidence intervals.
Six numbers, quarterly
Retrieval rate — the share of prompts triggering a web search. Currently 18 to 35 percent depending on measurement, and falling. If it keeps falling, the entity lever grows and the content lever shrinks.
Citation breadth — unique domains receiving referrals. Off its late-2025 peak but well above a year earlier. If it resumes narrowing, the specialist thesis is losing.
Ad density inside answers — around a quarter of ChatGPT responses.
Checkout share of agentic traffic — around 3 percent.
Platform share volatility — a twenty-point swing in the last twelve months.
First-party reporting — whether clicks and CTR ever arrive.
Each has a threshold at which your allocation should change. Write those thresholds down now, while nobody’s under pressure, and revisit them every quarter.
Chapter 16 — The Short Version
The temptation with a shift this size is to treat it as a new channel to conquer. A new acronym, a new tool category, a new agency, a new line item.
That reading is wrong, and it’s expensive.
Most of what gets sold under this heading is either ordinary excellence renamed, or speculation with no evidence behind it. Getting into the retrieval set dominates everything downstream. Seventy percent of AI Overview citations come from domains already on page one. Pages ranking first get cited three and a half times as often as pages beyond position twenty. When researchers built a benchmark to test AI-era content tactics head-to-head against traditional optimization, traditional optimization won. The cosmetic layer — restructuring, schema, chunk formatting — measures at or near zero in the best controlled experiments available.
What’s genuinely new is smaller and harder. Four things.
The entity layer, because a machine can’t judge the credibility of something it can’t identify, and because most questions never trigger a search at all — meaning what the model already believes about you is the whole game for most of them.
The corroboration layer, because how many independent documents agree is an input to a scoring function. Third-party agreement isn’t reputation management in the soft sense. It’s arithmetic.
The specificity layer, because one question becomes many, and because most of what people ask an assistant doesn’t appear in any keyword tool.
The agent layer, because a brand an assistant can’t read, price, or transact with is a brand it routes around.
And the thing all of it demands is a tolerance for uncertainty that most marketing organizations aren’t built for. You’re being asked to invest in a channel where impressions are visible but clicks aren’t, where the same measurement run twice disagrees with itself, and where the dominant effect is influence you can’t observe. The organizations that handle this well will be the ones that report honestly — rolling averages, separated metrics, explicit statements of what isn’t known — rather than the ones with the most confident dashboard.
There’s a version of this book that would have been more fun. It would have said the click is dead, the rules are rewritten, here are eleven tactics. Several of those tactics would have been the ones that measure at zero.
The real instruction is quieter:
Be a company a machine can name. Be reachable. Be backed up by people who don’t work for you. Be specific enough to answer the actual question. Be usable by an assistant acting for a customer.
Do those five things and you become the answer. Not because you gamed the system, but because when a system goes looking for a credible, consistent, specific, corroborated source, that is what it finds.
The woman in Ohio is going to ask her question tonight.
Somebody is going to be in that answer.
THE TOOLKIT
Appendix A — The Audit
One pass. Score each line pass, partial or fail. Anything graded A that fails is urgent. Lines phrased as prohibitions mean the opposite: a fail there means you’re doing something you should stop.
Entity
| Check | Grade |
|---|---|
| A canonical facts page exists, with positioning sign-off | A− |
| Homepage or About page states identity in plain indexable text | A− |
| Organization schema present, with a complete sameAs | A− |
| Wikidata item exists, is referenced, and its Q-ID is in sameAs | A− |
| LinkedIn, Crunchbase and directory descriptions match the canonical statement | A− |
| Press boilerplate on the last ten releases matches current positioning | A− |
| Named experts have author pages, consistent bylines and author markup | B |
| Knowledge Panel claimed, if one exists | B |
| A quarterly consistency review has a named owner | A− |
Reachability
| Check | Grade |
|---|---|
| robots.txt does not disallow OAI-SearchBot, PerplexityBot or Claude-SearchBot | A |
| Nobody has added named crawler groups that bypass the global rules | A |
| Nobody has blocked GPTBot believing it controls citation | A |
| Nobody has blocked Google-Extended believing it controls AI Overviews | prohibition |
| Cloudflare and WAF settings audited against the September 2026 defaults | A |
| Key content appears in view-source, and in a curl fetch using each bot’s user agent | A |
| AI-crawler 404 rate measured, worst offenders fixed | A |
| No noindex or nosnippet on pages intended for AI answers | A |
| Bot mitigation reviewed against agentic browser traffic | B+ |
| Search Console AI opt-out not enabled | A |
Corroboration
| Check | Grade |
|---|---|
| Third-party roundups and review pages cited for your top ten commercial queries are mapped | B+ |
| Factual errors about you in those sources have been corrected | B+ |
| Digital PR targets the outlets that actually get cited | B+ |
| At least one original research study published or commissioned annually | B |
| A named, disclosed presence exists on the two key community platforms | B |
| A sentiment-neutral review solicitation program is running | B |
| Nothing has ever been seeded, bought, or sentiment-incentivized | prohibition |
| The FTC review rule is in every agency contract | A |
| A YouTube program exists with real substance | B− |
Specificity
| Check | Grade |
|---|---|
| A question inventory built from real customer language exists | B+ |
| Attribute intersections are mapped, not just head terms | B |
| Questions scored for winnability: current citations, current rank, information gain | B+ |
| Prices and specifications published wherever commercially possible | B+ |
| Constrained-recommendation pages exist for the top intersections | B |
| Fair comparison content exists, including against competitors | B |
| Content is dated and on a defined refresh cycle | B |
| No mass-generated permutation pages | prohibition |
Usability
| Check | Grade |
|---|---|
| Product feed complete, refreshed on a defined cadence | A |
| Product Highlight and Product Detail populated across the catalogue | A |
| Q&A attributes populated in the feed | A |
| Manuals and spec sheets attached via document links | A |
| /.well-known/ucp published, where Universal Cart is live | A− |
| Price, availability and fulfilment terms in the initial HTML | A |
| Quarterly agent-completability test run on core journeys | B+ |
Measurement
| Check | Grade |
|---|---|
| Search Console generative AI reports and Bing AI Performance enabled | A |
| GA4 AI Assistant channel verified, with the Direct gap documented | A |
| AI-bot log analysis running | A |
| A question set of 120–200, versioned and frozen quarterly | B+ |
| One synthetic-prompt tool and one real-prompt tool in place | B |
| Only two-to-four-week rolling averages reported | A |
| Mention rate and citation rate reported separately | A |
| Monthly accuracy audit of answers about the brand | A |
| Controllable inputs reported alongside outputs | B+ |
| Survey questions capturing AI-assisted research added | B |
Appendix B — The Question Set
Shape
Target 120 to 200 questions.
| Segment | Share | Count at 150 |
|---|---|---|
| Constrained recommendation | 40% | 60 |
| Comparison | 20% | 30 |
| Brand-specific | 15% | 23 |
| Implementation and troubleshooting | 15% | 22 |
| Category education | 10% | 15 |
What to record for each
| Field | Note |
|---|---|
| Question text | Verbatim, in customer language |
| Segment | From the table above |
| Source | Sales call, support ticket, site search, forum, competitor position |
| Commercial value | High / medium / low |
| Current rank for the nearest search equivalent | From Search Console |
| Baseline cited domains | Recorded across at least seven runs |
| Baseline mention, yes or no | Averaged across runs |
| Information we hold that others don’t | Honest answer, or “none” |
| Winnable | Yes / no / not yet |
Rules
Source from harvested customer language, not from a brainstorm.
Include questions you currently lose. A set built from your strengths teaches you nothing.
Include your top three competitors’ strongest positions.
Cover the attribute matrix, not just head terms.
Freeze for a quarter. Version any change and treat the series as broken there.
Run each question at least seven times a day. Report rolling averages only.
Attribute matrix worksheet
For each core buying question, fill in your values and generate the intersections.
| Dimension | Your values |
|---|---|
| Company size or household type | |
| Industry or use context | |
| Integration or compatibility requirements | |
| Budget band | |
| Deployment or delivery model | |
| Compliance or regulatory regime | |
| Buyer sophistication | |
| Specific job to be done |
Three to five dimensions with three to five values each generates more intersections than you can serve. Rank by commercial value and honest winnability, then build the top twenty.
Appendix C — What the Evidence Says
Every tactic in this book, graded. Where studies conflict, both are named.
| Tactic | Evidence | Grade |
|---|---|---|
| Server-side render primary content | Vercel (1.3B crawler requests, 2024); searchVIU (23 crawlers, 69% cannot execute JS, 2025) | A |
| Fix AI-crawler 404s and redirect chains | Vercel: ~35% of ChatGPT and Claude fetches hit 404s | A |
| Correct crawler access control by user agent | First-party documentation from OpenAI, Anthropic, Perplexity, Google | A |
| Product feed completeness and the eight Merchant Center attributes | Google first-party specification, stated AI-surface purpose | A |
| Never seed or sentiment-incentivize reviews | FTC rule effective October 2024; enforcement live 2026 | A |
| Report rolling averages only | St. Gallen: run-to-run source overlap 0.34–0.42 | A |
| Entity consistency and Wikidata | Documented grounding mechanism; strong correlational support | A− |
| /.well-known/ucp manifest | Google first-party specification; Universal Cart live in the US | A− |
| Earned third-party mentions | Ahrefs: strongest correlate; Muck Rack: 82% of citations earned | B+ |
| Digital PR aimed at cited outlets | Muck Rack; third-party listicles ~80% of listicle citations | B+ |
| Publishing prices and specifications | Largest controlled experiment: price presence among the strongest factors, in a synthetic testbed | B+ |
| Agent-completability testing | Agentic traffic composition; 70% product and search routes | B+ |
| Organizing coverage around real question-form queries | Washington University: 6.8x activation on question-form queries (query form, not page formatting) | B |
| Original data, statistics and named sources | Academic GEO study; partially challenged by C-SEO Bench | B |
| Content recency and refresh cycles | Seer: 65% of AI bot hits target content under a year old | B |
| Passage self-containment | Documented chunking mechanism; direct effect unmeasured | B− |
| YouTube presence | Strongest single correlate (r ≈ 0.74); causality unestablished | B− |
| Structural rewriting for chunkability | Largest controlled experiment: negligible. C-SEO Bench: sometimes harmful | C |
| JSON-LD schema as a citation lever | Ahrefs difference-in-differences (n=1,885): −4.6% to +2.4%; Google: no special markup needed | C− |
| Mass-generated permutation pages | Content volume is the weakest correlate measured (r ≈ 0.18) | C− |
| llms.txt | Zero frontier-crawler fetches in a 900-domain, seven-month log study | D |
| Blocking Google-Extended for AI Overview control | Google documentation: does not affect Search inclusion or AI features | D |
| Keyword stuffing | −8%, the worst-performing method tested | D |
| Manipulating AI recommendations | Works, then degrades for everyone; plus regulatory exposure | F |
Appendix D — Glossary
AEO / GEO — Answer engine optimization, generative engine optimization. Used interchangeably. The practice of getting a brand included and well-positioned in AI-generated answers.
Chunk — A segment of a document, typically a few hundred tokens, indexed independently. The actual unit of retrieval.
Citation rate — Share of tracked questions where your domain appears as a source. Different from mention rate.
Entity — A specific, identifiable thing — a company, person, product — that a knowledge graph can resolve and attach facts to.
Ghost citation — An appearance where your site is cited as a source but your brand is never named in the answer. Roughly 62 percent of appearances.
Grounding — Checking generated claims against retrieved sources. Produces support scores, and can suppress low-confidence output entirely.
Information gain — What a passage offers that the passage it’s compared against does not. The one content principle that survives the evidence.
Mention rate — Share of tracked questions where your brand is named in the answer. The recommendation metric.
Query fan-out — Breaking one question into several sub-questions run in parallel. Confirmed by Google; the number of sub-questions is not public.
RAG — Retrieval-augmented generation. Retrieve relevant passages, put them in the model’s context, generate an answer grounded in them.
Retrieval threshold — The score below which a system answers from its own knowledge without searching. Google’s documented default is 0.3.
Share of voice — How often and how prominently a brand appears in AI answers relative to competitors. Widely sold, rarely sampled adequately.
UCP / ACP / MCP / A2A — Universal Commerce Protocol (Google), Agentic Commerce Protocol (OpenAI and Stripe), Model Context Protocol (Anthropic-originated), Agent2Agent (Linux Foundation).
Appendix E — Sources
Every figure in this book was verified against a primary source in August 2026. Platform behaviour, market share and regulation in this field move faster than publishing does, so check the date on anything that matters to a decision.
Research
Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande. “GEO: Generative Engine Optimization.” ACM SIGKDD 2024. https://arxiv.org/abs/2311.09735
Puerto, Gubri, Green, Oh and Yun. “C-SEO Bench: Does Conversational SEO Work?” NeurIPS 2025, Datasets and Benchmarks Track. https://arxiv.org/abs/2506.11097
Vishwakarma, Kumar and Jamidar. “What Gets Cited: Competitive GEO in AI Answer Engines.” SIGIR 2026. https://arxiv.org/html/2605.25517v1
Xu, Iqbal and Montgomery, Washington University in St. Louis. “Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact.” May 2026. https://arxiv.org/html/2605.14021
Schulte, Bleeker and Kaufmann, University of St. Gallen. “Don’t Measure Once: Measuring Visibility in AI Search.” April 2026. https://arxiv.org/pdf/2604.07585
Qin et al., Google Research. “Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting.” NAACL 2024. https://aclanthology.org/2024.findings-naacl.97/
Nestaas, Debenedetti and Tramèr, ETH Zürich. “Adversarial Search Engine Optimization for Large Language Models.” https://arxiv.org/html/2406.18382
Bliey and Chatwin. “Answer Engine Optimization: How Agentic AI Reshapes SEO.” Stanford GSBGEN 390, July 2026. SSRN.
Platform documentation
Google Search Central, AI features. https://developers.google.com/search/docs/appearance/ai-features
Google Search Central, common crawlers. https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers
Google Cloud, Vertex AI Search parsing and chunking. https://docs.cloud.google.com/generative-ai-app-builder/docs/parse-chunk-documents
Google Cloud, Check Grounding API. https://docs.cloud.google.com/generative-ai-app-builder/docs/check-grounding
Google Developers Blog, grounding with Google Search. https://developers.googleblog.com/en/gemini-api-and-ai-studio-now-offer-grounding-with-google-search/
Google Developers Blog, Universal Commerce Protocol. https://developers.googleblog.com/under-the-hood-universal-commerce-protocol-ucp/
OpenAI bots and commerce documentation. https://developers.openai.com/api/docs/bots
Anthropic crawler documentation. https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
Perplexity crawler documentation. https://docs.perplexity.ai/docs/resources/perplexity-crawlers
Microsoft Learn, chunking documents for retrieval. https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-chunk-documents
Cloudflare, “Content Independence Day: AI options,” July 2026. https://blog.cloudflare.com/content-independence-day-ai-options/
Google patents: generative summaries for search results (US11769017B1), search with stateful chat (US20240289407A1), pairwise ranking prompting (US20250124067A1).
Measurement and industry studies
Pew Research Center, “Google users are less likely to click on links when an AI summary appears,” July 2025. https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/
SparkToro, zero-click search analysis, June 2026. https://sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click/
Ahrefs, brand visibility correlations. https://ahrefs.com/blog/ai-brand-visibility-correlations
Ahrefs, schema and AI citations. https://ahrefs.com/blog/schema-ai-citations/
Ahrefs, most-cited domains in Gemini. https://ahrefs.com/blog/most-cited-domains-gemini/
Ahrefs, AI traffic volume study. https://ahrefs.com/blog/ai-traffic-increase/
Profound, AI platform citation patterns. https://www.tryprofound.com/blog/ai-platform-citation-patterns
Profound, how ChatGPT sources the web. https://www.tryprofound.com/blog/chatgpt-citation-sources
Semrush, ChatGPT search insights (clickstream). https://www.semrush.com/blog/chatgpt-search-insights/
Semrush and Kevin Indig, the ghost citations study. https://www.semrush.com/blog/the-ghost-citations-study/
Similarweb, AI search statistics 2026. https://aisearch.similarweb.com/blog/gen-ai-stats/
Similarweb, most-cited domains by LLMs. https://www.similarweb.com/blog/marketing/geo/most-cited-domains-llms/
BrightEdge, why AI engines cite different sources but recommend the same brands. https://www.brightedge.com/resources/weekly-ai-search-insights/ai-search-same-brands-different-sources
BrightEdge, YouTube presence in AI search. https://www.brightedge.com/resources/weekly-ai-search-insights/youtube-presence-ai-search
Muck Rack, GenerativePulse. https://media.muckrack.com/static/reports/2025/MuckRack-GenerativePulse2025-1.pdf
Seer Interactive, AI brand visibility and content recency. https://www.seerinteractive.com/insights/study-ai-brand-visibility-and-content-recency
Vercel, the rise of the AI crawler. https://vercel.com/blog/the-rise-of-the-ai-crawler
searchVIU, what AI systems actually see in schema markup. https://www.searchviu.com/en/schema-markup-and-ai-in-2025-what-chatgpt-claude-perplexity-gemini-really-see/
Microsoft Clarity, AI traffic conversion study. https://clarity.microsoft.com/blog/ai-traffic-converts-at-3x-the-rate-of-other-channels-study/
Brave, agentic browser security. https://brave.com/blog/comet-prompt-injection/
John Mueller on llms.txt, Search Engine Roundtable. https://www.seroundtable.com/google-ai-llms-txt-39607.html
Legal and regulatory
FTC, rule on the use of consumer reviews and testimonials. https://www.ftc.gov/business-guidance/resources/consumer-reviews-testimonials-rule-questions-answers
FTC, Operation AI Comply. https://www.ftc.gov/news-events/news/press-releases/2024/09/ftc-announces-crackdown-deceptive-ai-claims-schemes
EU AI Act Article 50 transparency obligations, in force August 2026.
Moffatt v. Air Canada, 2024 BCCRT 149.
Walters v. OpenAI, Georgia, 2025. Starbuck v. Google, filed October 2025.
Regional Court of Munich, case 26 O 869/26, May 2026.
Wikipedia:Conflict of interest. https://en.wikipedia.org/wiki/Wikipedia:Conflict_of_interest
Market and budget
McKinsey QuantumBlack, the agentic commerce opportunity. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-agentic-commerce-opportunity-how-ai-agents-are-ushering-in-a-new-era-for-consumers-and-merchants
Morgan Stanley Research, agentic commerce market impact outlook. https://www.morganstanley.com/insights/articles/agentic-commerce-market-impact-outlook
Gartner, 2026 CMO spend survey. https://www.gartner.com/en/newsroom/press-releases/2026-05-11-gartner-2026-cmo-spend-survey-finds-cmos-allocate-15-point-3-percent-of-marketing-budgets-to-ai-but-only-30-percent-are-ready-to-scale-ai-capabilities
NIQ, CMO outlook 2026. https://nielseniq.com/global/en/insights/report/2025/cmo-outlook-for-2026/
IPA, Binet and Field effectiveness research. https://ipa.co.uk/knowledge/effectiveness-research-analysis/les-binet-peter-field
Wikimedia Foundation, new user trends on Wikipedia. https://diff.wikimedia.org/2025/10/17/new-user-trends-on-wikipedia/
Put the book to work
The levers in this guide are the same ones we run for clients. Start with generative engine optimization, win the direct answer with answer engine optimization, fix retrieval with technical SEO, earn consensus through digital PR, and keep score with AI visibility monitoring. Our methodology maps chapter by chapter.
Read next on visibility.partners
- Generative Engine Optimization servicesChapters 5–8 delivered as a retained program.
- Answer Engine Optimization servicesWinning the direct answer and the sub-query.
- Technical SEO servicesThe retrievability work from Chapter 6.
- Digital PR servicesCorroboration and citation sources from Chapter 7.
- AI visibility auditAppendix A, run for your brand across five engines.
- AI visibility monitoringAppendix B prompt sets, re-run monthly.
- Our methodologyHow the book's 90-day plan maps to an engagement.
- What is answer engine optimization?A short primer on the ideas in Chapters 2 and 8.
Primary sources and further reading
Every claim in this book is grounded in vendor documentation, published research, or our own measurement. These are the primary sources worth reading directly.
- OpenAI platform documentation — How ChatGPT retrieval, browsing and tool use are documented by OpenAI.
- OpenAI: search in ChatGPT — OpenAI's own description of search behaviour and publisher citation.
- Google Search Central: AI features and your website — Google's guidance on AI Overviews, AI Mode and eligibility.
- Google: creating helpful, reliable, people-first content — The quality guidance underpinning E-E-A-T signals.
- Google structured data reference — Supported schema types for rich results and machine parsing.
- Schema.org vocabulary — The entity vocabulary used throughout Chapters 5 and 6.
- Perplexity developer documentation — How Perplexity retrieves, ranks and attributes sources.
- Google Gemini API documentation — Grounding, search retrieval and citation behaviour in Gemini.
- Anthropic documentation — Claude's retrieval, citations and tool-use model behaviour.
- Bing webmaster guidelines — Indexing and quality guidance behind Copilot's web answers.
- GEO: Generative Engine Optimization (Aggarwal et al., 2023) — The peer-reviewed paper that named generative engine optimization.
- Retrieval-Augmented Generation (Lewis et al., 2020) — The original RAG paper behind the retrieval pipeline in Chapter 2.
- web.dev: Core Web Vitals — Performance thresholds referenced in the retrievability audit.
- The robots.txt specification — Baseline for AI crawler access decisions in Chapter 6.
Ranking is no longer enough
You need to be cited, mentioned, and recommended.
Get a free AI Visibility Report — see exactly where your brand appears across ChatGPT, Google AI, Gemini, Perplexity, and Copilot, and where competitors are winning instead.