LUCID: Universal Auditing of Distilled Large Language Models
Abstract
The growing transparency of large language models (LLMs) makes distillation into smaller models an inevitable practice, allowing users to cheaply inherit advanced capabilities such as reasoning. Yet this trend also exposes model providers to new risks: unauthorized data distillation may misappropriate the teacher model’s valuable functions, resulting in copyright violations, privacy leaks, and other serious harms. Existing fingerprinting techniques mainly focus on detecting complete model theft, offering little protection for specific functional capabilities, and many require white-box access, limiting real-world applicability. In this work, we propose LUCID (LLM distillation Unveiled via invarianCe auditor Infringement Detection), the first auditing framework designed for scenarios where the suspect model’s internals are inaccessible, tailored to identifying the misappropriation of a victim model’s specific capability. By utilizing self-reflective prompts to elicit internal judgments, LUCID extracts invariant judge-token representations to distinguish infringing models from non-infringing ones. Theoretical analysis substantiates the generalization ability of our approach, while empirical results demonstrate its effectiveness in reliably identifying unauthorized distillation across diverse model architectures and deployment settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.