PART THREE — RUNNING IT

Chapter 10 — What Actually Works

Most research in this field is produced by companies selling tools in it. This chapter grades every tactic: what the evidence supports, what is genuinely contested between studies that disagree, and what is being sold to you with no measured effect behind it — so budget decisions rest on graded evidence.

From Becoming the Answer by Jeremy Osborn · 1,055 words

This field has a research problem. Not a shortage — the volume is impressive — but most of it is produced by companies selling tools in the category, most of it is correlational, and the studies that disagree with each other rarely get read together.

You are going to be asked to approve budget against this. So here is the scoreboard, sorted into what we know, what’s contested, and what’s being sold to you.

Everything below is graded:

GradeMeaning
AMultiple independent measurements, or a first-party platform specification. Act on it.
BOne good study or consistent practitioner measurement with a plausible mechanism. Act, verify locally.
CContested. Correlational, confounded, or the studies disagree. Do it if it’s cheap.
DNo evidence, or evidence against. Don’t fund it.
FEvidence of harm, or unlawful. Prohibit it.

A plus or minus modifies within a grade.

What we know

Server-side render anything you want cited. Grade A. Measured independently by two organizations with no dissenting result. Most dedicated AI crawlers cannot execute JavaScript. Content that only exists after hydration doesn’t exist.

Fix your 404s and redirect chains. Grade A. Roughly a third of AI crawler fetches on the average site hit dead URLs. Mechanically obvious, cheap, and directly increases how much of your real content gets retrieved.

Get your crawler configuration right. Grade A. First-party documentation from every platform. The bots that gate training are not the bots that gate citation, and getting it backwards is common and invisible.

Complete your product feed. Grade A. Google and OpenAI both publish specifications. The eight Merchant Center attributes are explicitly connected to AI surfaces by Google’s own documentation. This is the least ambiguous work in the book.

Establish your entity. Grade A−. Low cost, documented mechanism, strong correlational support. Nobody has isolated the effect of fixing a specific inconsistency, which is why it isn’t a straight A.

Earn third-party mentions. Grade B+. The strongest correlate in the largest correlational study, and separately, corroboration count is a documented input to grounding. Two independent reasons pointing the same way, neither of which proves the other.

Publish real prices, specifications and original data. Grade B+. Price presence carried one of the largest effects in the biggest controlled experiment, with the caveat that the experiment ran in a synthetic testbed rather than in Google.

Organize content around real question-shaped queries. Grade B. Question-form queries trigger AI Overviews nearly seven times more often than other query forms. Note carefully: that’s about how users ask, not about how you format your headings.

Refresh content on a cycle. Grade B. Around 65 percent of AI crawler attention goes to content published within the past year. Recency is a retrieval factor.

Invest in YouTube. Grade B−. The strongest single correlate in the brand-visibility data, and the most-cited domain inside Google AI Overviews. Causality unestablished, confound obvious, cost of testing low.

What’s contested

Restructuring content for “chunkability.” Grade C. The largest controlled experiment found formatting and structure effects minimal across every model tested. A peer-reviewed benchmark found conversational rewrites frequently harmed ranking. Write well because it’s good writing. Don’t fund a reformatting project.

Schema markup as a citation lever. Grade C−. This is the most confidently asserted tactic in the field, so the evidence deserves stating plainly.

Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 matched control pages over eight months, using difference-in-differences. The effect on Google AI Overviews was −4.6 percent, small but statistically significant. On AI Mode, +2.4 percent. On ChatGPT, +2.2 percent. Both of the last two are indistinguishable from zero.

A separate mechanical test put eight product prices on a page, distributed across visible HTML, JavaScript-rendered content, JSON-LD only, and hidden markup, then asked each system to find them. No system found prices that existed only in JSON-LD.

And Google’s own documentation says, in as many words, that there’s no special structured data you need to add to appear in these features.

This does not mean abandon schema. It means stop funding it as a citation lever. Structured data still does real work for entity disambiguation, for Bing and Copilot, for classic rich results, and for commerce feeds. “Getting cited in AI answers” isn’t one of its jobs.

“Share of voice” as a stable metric. Grade C. Non-determinism, personalization and arbitrary prompt-set design are unpriced in every vendor’s reporting. Chapter 11 deals with this.

What’s being sold to you

llms.txt. Grade D. Google’s John Mueller, on the record: “no AI system currently uses llms.txt.” A server-log study across roughly 900 domains over seven months recorded 1,227 requests for the file — fewer than seven per day across all sites — of which two-thirds came from a single data broker and a third from Chrome browsers, meaning humans. Zero came from a frontier AI lab’s crawler.

There is one genuine signal: Chrome added an llms.txt audit in 2026, filed deliberately under agentic browsing rather than SEO. That’s where the file may eventually matter — agent tooling, not search visibility.

Publishing one costs an hour and harms nothing. Don’t let it onto a roadmap as a visibility driver, and don’t pay an agency for it.

Blocking Google-Extended to control AI Overviews. Grade D. It does not do that. Google-Extended governs whether content may be used to train and ground Gemini products. AI Overviews and AI Mode are features inside Google Search, which the token does not touch.

Keyword stuffing. Grade D. Measured at −8 percent — the worst-performing tactic in the academic study of content optimization. It doesn’t fail to help. It hurts.

Manipulating recommendations. Grade F. Covered in Chapter 13. It works, briefly, and then it stops working for everyone including you.

The through-line

Read together, the strongest studies say the same thing, and it isn’t what the category sells.

Getting into the retrieval set dominates everything downstream. Seventy percent of AI Overview citations come from domains already on page one. Pages ranking first get cited three and a half times as often as pages beyond position twenty. Traditional search optimization beat purpose-built conversational tactics in the one peer-reviewed head-to-head. Relevance, recency and candidate position dwarfed every cosmetic content factor in the largest controlled experiment.

What’s genuinely additive is the entity work, the corroboration work, the specificity work, and the agent-readiness work. That’s where the incremental return lives, and it’s four things rather than forty.

Back to the full contents of Becoming the Answer