How AI Search Works

How Do AI Search Engines Choose Sources?

How AI search engines choose sources. The retrieval, ranking, and citation signals that decide whether your brand appears in ChatGPT, Gemini, and Perplexity answers.

July 19, 202611 min readBy Jeremy Osborn

AI search engines don't pick sources at random. Every citation is the output of a pipeline: retrieval, reranking, and citation each apply distinct signals. Understanding that pipeline is the difference between guessing at AI SEO and engineering it.

  • AI citations
  • AI search sources
  • how LLMs cite

Stage 1: retrieval

Retrieval is a recall problem — pulling a broad candidate set the engine could use to answer the query. It's optimized for coverage over precision.

The signals that matter here are boring but decisive: can the engine's crawler reach your pages, does your site render cleanly, is your information architecture legible enough for content to be clustered, and are your pages present in whichever embedding index the engine uses. Domains that fail at retrieval never enter the pipeline — no amount of downstream work fixes that.

Stage 2: reranking

Reranking is a precision problem. Given the candidate set, the engine scores each passage for how well it answers the query and how confidently it can be lifted into an answer.

Signals: passage-level clarity, topical fit to the specific query, freshness, and structural cues (question-first headings, tight answer blocks, structured lists). This is the stage where classic AEO work pays off — the same content patterns that win featured snippets tend to win reranker attention.

Stage 3: citation

Citation is a trust problem. The engine has to decide which of the reranked candidates to name — and which to lift verbatim into the answer.

This is where entity strength and third-party authority do the heaviest lifting. A page from a confirmed, well-cited domain will win the citation slot over a technically equivalent page from an unknown domain almost every time. It's also where sentiment can flip a placement: mentions inside negative context can be suppressed by the engine's safety and quality layers.

What signals matter most

  • Clear entity identity across the open web
  • Depth and consistency of topical coverage
  • Third-party citations on trusted domains in your category
  • Structured, extractable passages
  • Recency of publication and updates
  • Absence of contradiction across your own content

How engines differ

The three stages exist in every major engine, but the weightings differ. Perplexity leans hardest on third-party citation trust — its whole value prop is source attribution. ChatGPT balances entity strength (from training) with live retrieval trust. Google AI Overviews inherit the classic search authority stack and add answer-format signals on top. Gemini emphasizes grounding sources and Google-ecosystem confirmation.

A program that treats all four the same wins in one and loses in another. Segmenting reporting and tactics by engine is how mature programs avoid that trap.

Frequently asked

Can I influence AI citations directly?

Yes — through content, entity, and third-party citation work. Our AI SEO agency runs the full stack. The signals are engineered, not accidental.

Do LLMs use Google's index?

Some use Google or Bing under the hood; others use their own crawlers and third-party APIs. The specifics vary by engine and change over time.

Why is my competitor cited more often?

Usually stronger entity signals, denser third-party authority, or more extractable content — sometimes all three. An AI visibility audit isolates which.

Keep reading

Ranking is no longer enough

You need to be cited, mentioned, and recommended.

Get a free AI Visibility Report — see exactly where your brand appears across ChatGPT, Google AI, Gemini, Perplexity, and Copilot, and where competitors are winning instead.