How does AI search choose which sources to cite?

Learning center4 min read

An AI search engine chooses its sources in three stages. A conventional search index returns candidate pages for the queries the model wrote. A ranking step splits those pages into passages and scores each passage against the question. The language model then writes the answer from the top passages and attaches citations to the ones it used. A page can drop out at any of the three stages, and most pages drop out at the first.

The exact scoring is not published by any vendor, and it differs between ChatGPT search, Perplexity, Google AI Overviews and the others. What they share is the pipeline, and the pipeline is what makes some pages reliably citable and others reliably ignored.

Stage one: the search index

The model does not search the whole web. It sends queries to an index: Bing's index in the case of ChatGPT search and Copilot, Google's own for AI Overviews and AI Mode, and a mix of own and licensed indexes for Perplexity. Only pages in that index can be retrieved.

The queries themselves are written by the model. A question like "which running shoes are good for flat feet" may become several queries about stability shoes, arch support and specific brands, a behavior known as query fan-out. Each query returns its own candidates, so a page can be reached through a sub-question the user never typed.

At this stage the usual search signals apply: relevance to the query, the authority of the site, freshness, and whether the page is indexed at all. A page that ranks on page three for a query rarely becomes a candidate. A page blocked to the vendor's crawler is never a candidate.

Stage two: passage selection

The candidate pages are fetched and cut into passages. A passage is a paragraph, a list, a table or a short section. The system scores each passage against the question using a mix of keyword overlap and semantic similarity, and often runs a reranker over the best candidates.

The scoring rewards passages that answer the question by themselves. In practice that means:

  • A direct answer in the first sentence. Setup can come after it.
  • One idea per passage. A paragraph that covers four points matches four questions weakly instead of one question strongly.
  • Concrete details: a number, a date, a name, a step. Vague passages score low because they look like any other vague passage.
  • Plain structure. Lists and tables are easy to extract; text inside interactive widgets or loaded only by JavaScript may never be seen.

A passage also competes against passages from other pages saying the same thing. If ten pages give the same answer, the ones from sites the system already trusts tend to win, and the rest are redundant.

Stage three: writing and citing

The model receives the selected passages and the question and writes the answer. Citations are attached while it writes: when a sentence is drawn from a passage, the passage's URL is linked to it. A passage that reached the model but was not used gets no citation, so being retrieved is not the same as being cited.

Which passages get used depends on how the model composes the answer. Passages that contain the kind of statement the answer needs, such as a definition for a "what is" question or a list of options for a "best" question, are the ones that end up quoted. A page that is on topic but phrased as a story or an opinion is often read and then left out.

This stage also decides which brands are named. For a "best tools for X" question the model lists the names that appear in the passages it read. A brand that is absent from the comparison pages, review sites and forums the retriever returned does not appear in the answer, whatever its own website says.

Why do citations concentrate on a few domains?

Across many questions, a small group of domains receives most of the citations. Wikipedia, Reddit, YouTube, major news sites, large retailers and well-known review sites recur in almost every study of AI answers, including the AI search statistics Searcherries collected.

The pipeline explains the pattern. Those sites are in every index, rank well for many queries, have pages structured as direct answers, and are trusted enough that the ranking step prefers them when several pages say the same thing. Each stage filters toward them, and the effect compounds.

The practical consequence is that presence on those domains counts as much as presence on your own site. A product that is discussed in a Reddit thread, listed on a review site and described in a comparison article has three more routes into the answer than one that is only described on its own home page.

What can a page do to be selected more often?

Each stage suggests a change.

For the index, make sure the page is crawlable by the relevant bots, is indexed by Google and Bing, and ranks for the sub-questions the model is likely to write, not only for the head term.

For passage selection, open each section with the answer, keep sections to one topic, use headings that read like questions, and put facts in lists and tables rather than prose.

For citation, write the kind of statement the question calls for. A definition page should define in the first paragraph. A comparison page should name the options and the criteria. A pricing page should state the price.

Whether any of this is working is a question of measurement. AI visibility tracking runs the same questions repeatedly across platforms and records which pages and brands are cited, which shows the pipeline's output directly.

Start a 14-day free trial
and get your AI visibility report