PART TWO — THE FIVE THINGS THAT DECIDE IT

Chapter 7 — Be Backed Up

Being talked about tracks AI visibility far more strongly than being linked to. This chapter shows where citations actually come from — Wikipedia, Reddit, YouTube, earned journalism, third-party listicles — how little the engines overlap with each other, what ghost citations are, and what genuinely earns third-party corroboration.

From Becoming the Answer by Jeremy Osborn · 1,672 words

Chapter 3 established the mechanism: how many independent documents verify a claim is an input to whether it survives. Chapter 5 established the correlation: being talked about tracks AI visibility far more strongly than being linked to.

These are two separate arguments, and neither proves the other. Both point at the same work — which is the most expensive lever, the slowest, the least controllable, and the one with the highest ceiling.

Where answers actually come from

The distribution of sources is not what most people assume, and it varies enormously by platform.

Wikipedia is the most-cited single domain in ChatGPT, taking somewhere between 5 and 13 percent of citations depending on the study and period. Reddit is first or second nearly everywhere: Ahrefs found it the single most-cited domain in Gemini at 29 percent of top-fifty mention share, with YouTube second. In Perplexity, YouTube leads at 31 percent, Reddit second.

But concentration is lower than the headlines suggest. Profound’s analysis of roughly 730,000 real ChatGPT conversations found the top ten domains accounting for only 12 percent of all citations. The tail is long.

Now the finding that should change how you allocate. BrightEdge compared five engines across nine industries and measured overlap two ways:

Which sources they cite: 16 to 59 percent overlap.

Which brands they recommend: 36 to 55 percent overlap.

The engines disagree substantially about where to pull information from. They agree much more consistently about which brands belong in the answer.

Kevin Indig found the same thing from another direction: 91 percent of citations appear on only one of ChatGPT, Perplexity, or AI Overviews.

The strategic instruction falls straight out of that. Chasing per-platform citation tactics is a treadmill. Brand-level outcomes converge. Optimize the entity and the corroboration, not the platform.

Being cited is not being recommended

Bring this one to the first meeting where somebody shows you an AI visibility dashboard.

Semrush and Kevin Indig examined 3,981 domain appearances across 115 prompts, fourteen countries and four engines. Sixty-two percent were “ghost citations” — the site linked as a source, the brand never named in the answer. Only 13 percent got both.

The platform split is nearly inverse. ChatGPT cites your domain 87 percent of the time and names your brand 21 percent. Gemini cites 21 percent and names 84 percent.

These are two different outcomes and they need two different metrics. Being the footnote is not being the recommendation. If your dashboard reports one number called “visibility,” find out which one it’s measuring.

What earns citations

The largest study available on this comes from Muck Rack, which analyzed over a million links from AI responses across ChatGPT, Claude and Gemini over six months. They sell PR software, so read it as interested research — but it’s the best-documented dataset in the category.

Earned media accounted for 82 percent of all citations (revised upward to 84 percent in a later update). Journalistic sources took roughly a quarter of all links. Half of all citations went to content published within the previous eleven months, and about 4 percent to content published in the previous week.

Seer Interactive triangulates the recency finding from server logs: around 65 percent of AI bot hits target content published in the past year, 89 percent within three years, only 6 percent older than six.

Recency is a retrieval factor, and it’s one of the few on this list you control directly. A refresh program that genuinely updates and re-dates evergreen content is doing retrieval work, not housekeeping.

Where commercial queries are decided

For “best X” and “X versus Y” — the queries nearest to revenue — the evidence is consistent and unwelcome.

Listicles are the largest single page type by citation share, around 22 percent. Of those, roughly 80 percent are third-party listicles. Only 20 percent are brand-authored. Comparison pages, counterintuitively, take under 3 percent; the engines appear to prefer pre-consolidated recommendations to side-by-side tables.

The concentration is severe: about thirty domains capture two-thirds of citations within a topic, and in product-comparison topics the top ten capture nearly half.

In B2B software specifically, an analysis of 57 million citations across fifty companies found Reddit at 28 to 31 percent, YouTube at 15 to 20, LinkedIn at 8 to 15 — and G2 at 11 percent on branded queries but out of the top fifteen entirely on unbranded ones. Brand-owned content was a small fraction of everything cited.

That G2 result is worth sitting with. Review sites show up when someone names your brand and vanish when they don’t. That’s a reputation-defense job, not a discovery job, and they deserve different budgets.

The practical implication: for commercial queries, the highest-leverage work isn’t on your website at all. It’s being accurately represented in the third-party roundups and category comparisons the engines actually cite. That is a PR and analyst-relations program, not a content program.

The most underrated surface here

Two independent measurements point at YouTube, and almost nobody’s strategy reflects it.

BrightEdge found YouTube averaging around 20 percent citation share across AI platforms — and 29.5 percent inside Google AI Overviews, where it is the single most-cited domain. Vimeo and TikTok sit at 0.1 percent each. This isn’t “video.” It’s YouTube.

Ahrefs found YouTube mentions the strongest single correlate of AI visibility across every platform they tested, higher than brand mentions, far higher than backlinks.

The two datasets disagree sharply on Perplexity — BrightEdge puts YouTube at under 10 percent there, Ahrefs at 31 — which is a reminder that these are different measurements with different denominators, not one finding confirmed twice. Causality is unestablished and the confound is obvious.

Still: the platform is owned by the company running the largest answer engine, the correlation is the strongest in the dataset, and the cost of testing is low. Fund an experiment.

The line you must not cross

Community platforms matter, which makes the temptation to manufacture presence on them large. Don’t.

The legal position is no longer ambiguous. The FTC’s rule on consumer reviews and testimonials took effect in October 2024 and prohibits, among other things: fake reviews including AI-generated ones; reviews compensated on condition of sentiment; undisclosed reviews by employees or their immediate relatives; company-controlled sites posing as independent; suppressing negative reviews through threats; and buying fake engagement.

Civil penalties reach $53,088 per knowing violation, and “per violation” can mean per review. Enforcement is live: a 2026 settlement over employees posing as users and incentivized five-star reviews produced a $4 million judgment.

Where the line actually sits is more nuanced than the panic suggests. Paying for reviews isn’t automatically illegal — conditioning payment on sentiment is. “Tell us how much you loved your visit and get a $5 coupon” violates the rule. “Tell us about your visit and get a $5 coupon” doesn’t. General solicitations to customers are safe even if employees respond, as long as no sentiment requirement is attached. Employees may review if they clearly disclose the relationship.

The reputational position is worse than the legal one. A marketing agency publicly boasted in 2025 about running forty-plus fake Reddit accounts staged as authentic player discoveries; the post was deleted inside a day and both the agency and its client issued apologies. When researchers ran undisclosed AI-generated personas on Reddit for months, Reddit’s chief legal officer called it “deeply wrong on both a moral and legal level,” banned the accounts and issued formal legal demands to the university involved. The paper was withdrawn.

If Reddit will pursue a university, it will pursue you.

The asymmetry is the whole argument. The upside of manufactured community presence is a temporary bump in a metric nobody can measure reliably. The downside is a statutory penalty per item, a permanent negative record that AI systems will retrieve for years, and a press cycle. There is no version of this trade that works.

What legitimate participation looks like

In descending order of defensibility:

Answer questions in your own name, with your affiliation visible. Reddit’s policy targets repeated unsolicited promotion, not participation. A named employee giving a genuinely useful answer is within the rules and within the norms.

Host AMAs through official channels, coordinated with moderators.

Publish on YouTube. Given everything above, this is the highest-leverage community surface available and it carries no astroturfing risk at all.

Solicit reviews generally, without sentiment conditions, from actual customers.

Fix the substance. Product and service problems become AI-visible problems, because the community discussion becomes the retrieval corpus. No communications strategy outruns this.

The consensus test

The deepest point in this chapter is structural.

A system seeking corroboration finds three narratives about you. Owned — your website, blog, releases. One voice. Earned — press, analysts, third-party roundups. A second. Community — Reddit, reviews, forums, YouTube. A third.

If owned and earned claim excellence while the community reflects frustration, the system synthesizes toward the community signal. Not because it has taste, but because its architecture favors multi-source agreement over single-source claims.

The brands that win are the ones where the story on the website, the story told by third parties, and the story told by actual customers all point the same way.

That is not a marketing problem. It is an operating one, which is why this lever belongs to the executive team rather than the content calendar.

The corroboration checklist

Map the third-party roundups and review pages cited for your top ten commercial queries — 1 week

Correct factual errors about you in them — the fastest win available — 2–4 weeks

Build analyst and journalist relationships with the outlets that actually get cited, not the prestigious ones — ongoing

Publish original research nobody else has; make the data quotable — quarterly

Fund a YouTube program with real substance — ongoing

Put content on a refresh cycle; recency is retrieval — ongoing

Establish a named, disclosed presence on the two forums where your category is discussed — ongoing

Run a sentiment-neutral review solicitation program — ongoing

Put the FTC review rule in every agency contract — 1 day

Never buy, seed, or sentiment-incentivize anything

Back to the full contents of Becoming the Answer