RAFP: Identifying LLM Lineages via Rare-Region Fingerprints
Abstract
Large language models (LLMs) are increasingly adapted and redistributed through finetuning, parameter-efficient tuning, and quantization, making it difficult to verify whether a suspect model is derived from a particular pretrained model. Existing fingerprinting approaches often rely on modifying the protected model or construct behavioral signatures without explaining why they should persist under downstream adaptation. We introduce RAFP, a non-invasive framework for identifying LLM lineages through rare-region fingerprints. Our key insight is that downstream adaptation is concentrated on common high-density language behaviors, while low-probability prompt regions receive weak optimization signal and exhibit limited gradient alignment under finetuned distribution. We show theoretically that, under mild assumptions on downstream coverage and gradient alignment, the likelihood drift of behaviors in these rare regions remains bounded during finetuning. Building on this insight, RAFP searches for rare prompt–response pairs whose behaviors are shared across models in the same lineage but remain discriminative from unrelated models, without modifying model parameters. The resulting fingerprints support black-box ownership verification and can be generated post hoc for pre-trained models. Experiments across four LLM families and multiple downstream adaptations, including supervised finetuning, LoRA, quantization, prompt-template variation, and decoding changes, show that RAFP achieves strong fingerprint persistence and substantially outperforms prior fingerprinting baselines in black-box settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.