PART THREE — RUNNING IT

Chapter 11 — Measuring Without Fooling Yourself

You are being asked to manage a channel whose cost is measurable and whose benefit is not. This chapter replaces rank with citation rate, mention rate and recommendation rate, covers the free data sources most teams ignore, explains how to design and lock a question set, and what to actually report.

From Becoming the Answer by Jeremy Osborn · 1,448 words

Before any dashboard, tool selection or board slide, somebody in your organization has to say this sentence out loud:

We are being asked to manage a channel whose cost is measurable and whose benefit is not.

Search Console shows AI impressions but not clicks, click-through rate, queries or position. Bing shows citations but not traffic value. Analytics catches some referrals and loses the rest to Direct. And the dominant effect — being recommended in an answer nobody clicks — leaves no trace at all.

Any measurement plan that promises to close that gap is overpromising. What follows closes what can be closed and is honest about the rest, which is the achievable goal.

Rank is gone, and what replaces it

Under classic search there was one results page and one position. Under personalized, non-deterministic generation there is no universal ranking. A brand that appears first for one person may not appear at all for another.

What actually breaks rank as a metric isn’t the averaging. It’s the variance — across people, because of personalization, and across runs of the same question, because of stochasticity. A single observed position tells you almost nothing about either.

So the thing you’re estimating is your average position across the population of people who might ask. You can’t observe that population, so you approximate it with three things: a question set that represents real demand, enough repeated runs to average out randomness, and enough platform coverage to reflect where your buyers actually are.

Miss any of the three and you’re producing noise with invisible error bars.

The sampling requirement nobody meets

Researchers at the University of St. Gallen tested something the vendor category would rather nobody tested: run the same prompt over and over and see how much the answer moves.

Identical prompts on identical days produced source overlap of 0.34 to 0.42 and brand-mention overlap of 0.45 to 0.59. The instability held inside 24-hour windows, so it’s the model, not the news.

Their recommendation: at least seven runs per prompt per day for brand tracking to get the standard error below 0.10, at least eight for source coverage, and rolling windows of two to four weeks. Their conclusion, in their words: single observations are misleading.

Most commercial tools sample far below this, and almost none publish confidence intervals. Three rules follow, and they should be non-negotiable:

Never report a single-day figure. Rolling averages over two to four weeks only.

Ask every vendor in procurement how many runs per prompt per day they execute, and whether they publish confidence intervals. The answer is diagnostic. Most won’t have one.

Treat any before-and-after case study without disclosed repeated sampling as unproven. Including your own. Especially your own.

Two kinds of tool, and why they disagree

Every vendor sells “share of voice in AI answers.” They do not measure the same thing.

Synthetic prompting. The tool fires a prompt list you configure at model APIs or chat interfaces on a schedule, parses the responses, aggregates. Most of the category works this way. You control the question set and get per-prompt diagnostics. The weakness is fundamental: you’re measuring answers to questions you invented.

Real-prompt corpora. A smaller group claims to derive prompts from observed behavior rather than invention — Ahrefs states an index of over 473 million monthly prompts; Profound claims an index built from over 1.5 billion real conversations. Neither publishes an auditable methodology for how they obtain them at that scale, so treat the volume claims as directional rather than audited.

A working approach: buy one of each. They will disagree, and the disagreement is the methodology rather than a bug. Seeing both keeps you honest.

Rough entry pricing, useful mainly for scale: Profound from around $99 a month, Ahrefs Brand Radar from $199, Otterly from $29, Semrush’s AI toolkit around $99 per domain, Scrunch from $250, with Conductor and BrightEdge sold as enterprise. Verify before quoting — these move.

The structural point is that a serious measurement program costs a few thousand dollars a year, not a few hundred thousand. The constraint is analytical discipline, not budget.

The free sources you’re probably ignoring

Google Search Console generative AI reports, launched mid-2026. Impressions from AI Overviews and AI Mode, by page, country, device and date. No clicks, CTR, queries, position or conversions. One trap: impression counts differ by aggregation level, so property-level and page-level totals aren’t comparable.

Bing Webmaster Tools AI Performance report. Total citations, unique URLs cited daily, sample triggering queries, per-page counts. Free, and the only first-party citation count anybody publishes.

GA4’s AI Assistant channel. Sets medium to ai-assistant when it recognizes a referrer. Remember it’s a floor: in-app and mobile traffic often arrives with no referrer at all, and one study found 22 percent of AI Overview traffic misfiled as Direct.

Your server and CDN logs. The most underused source in the stack. Filter by AI user agent and you get crawl frequency by bot, which content is being fetched, how much budget is burning on 404s, and whether your robots.txt and firewall changes actually took effect. Free, and it answers questions no dashboard will.

Designing a question set worth measuring

In synthetic prompting, the question set is the entire ballgame. A tool tracking 25 prompts is a tripwire, not a measurement system.

For a mid-sized program, aim for 120 to 200 questions:

SegmentShare
Constrained recommendation — “best X for [situation]”40%
Comparison — “X vs Y”, “alternatives to Y”20%
Brand-specific — “is [you] good for X”15%
Implementation and troubleshooting15%
Category education10%

Four rules. Source them from the harvest in Chapter 8, not from a brainstorm. Include questions you currently lose — a set built only from your strengths produces a dashboard that flatters you and teaches you nothing. Include your top three competitors’ strongest positions. And freeze the set for a quarter; when you do change it, version it and treat the series as broken at that point.

What to report

Four numbers, each answering a distinct question. Don’t collapse them into a single “AI visibility score” — those scores hide exactly the trade-offs that matter.

Mention rate. Share of tracked questions where your brand is named in the answer, on a rolling basis. This is the recommendation metric.

Citation rate. Share where your domain appears as a source. Remember: 62 percent of appearances are citations without mentions, and ChatGPT and Gemini are nearly inverse on this. Reporting one and calling it the other is the most common error in the field.

Accuracy. What’s being said, and whether it’s true. This needs a person reading a sample of answers, not a sentiment classifier. It catches problems before they become Chapter 13 problems.

Competitive share. Your mention rate against your named competitors on the same question set. In a comparative system the relative number matters more than the absolute one.

Alongside those, report the inputs you control — entity consistency, share of key content server-rendered, AI crawler 404 rate, third-party roundups where you appear correctly, feed completeness. Inputs move before outputs do, and they’re the only part of the system where cause and effect are legible.

The ROI conversation

You will be asked what this returns. Here’s how to answer without lying.

What you can attribute: AI referral sessions and their conversion behavior, understanding the number is a floor. If that traffic converts materially better than other channels — and the evidence suggests it currently does — that’s a real, defensible, small number.

What you can’t attribute: the influence of appearing in answers nobody clicks. That’s most of the value, and there is no honest way to measure it directly today.

What to do about the gap. Three things.

Proxy it with competitive mention rate. If you’re named in 40 percent of relevant answers and your closest competitor is named in 15, that’s a share-of-consideration position with real economic content even if you can’t price it.

Instrument the qualitative side. Add “did you research this with an AI assistant?” to your demo request and post-purchase surveys. Self-reported attribution is imperfect and it’s currently the only line of sight into zero-click influence. Track it as a percentage over time.

And be explicit about the asymmetry in the business case. The entity and reachability work is small and largely one-time. The corroboration work is real money, but it’s mostly spend you’re already making on PR and content, redirected. You aren’t asking for a new budget line proportional to an unmeasurable return. You’re asking to reallocate an existing one toward mechanisms that are documented.

That last framing is what gets this funded. A request for new money against an unmeasurable return fails. A request to redirect existing money toward better-evidenced mechanisms succeeds.

Back to the full contents of Becoming the Answer