Cambium: Pruning-Robust Black-Box Fingerprinting for Large Language Models
Abstract
Large language models (LLMs) are powering a wide range of real-world applications. Making them valuable intellectual property requires substantial resources to train. Therefore, fingerprints are often embedded into LLMs to verify model ownership. In local deployment, LLMs are commonly fine-pruned to reduce storage cost and accelerate inference, especially in resource-constrained settings such as edge devices and small or medium-sized organizations. However, fine-pruning can erase embedded fingerprints and make model ownership unverifiable. We propose Cambium, a pruning-robust black-box fingerprinting method for LLMs. Our key observation is that a fingerprint is more likely to survive fine-pruning when its gradient is more consistent with the task gradient. To encourage this gradient consistency, Cambium first generates a verifiable and task-compatible fingerprint pair by selecting a target response supported by the model's natural output distribution. It then embeds the fingerprint through gradient-modulated supervised fine-tuning, which encourages consistency between the fingerprint gradient and the model task gradient. Comprehensive experiments with four LLMs on two different task datasets show that Cambium effectively embeds fingerprints, achieving a 100% fingerprint success rate (FSR) while preserving model utility. Across five increasingly aggressive fine-pruning settings, Cambium maintains an FSR of 71.8%–100%, while the three baselines reach 0% FSR in 46.7% of the evaluated settings. To verify ownership, Cambium achieves a 100% verification success rate (VSR) across all fine-pruning settings with at most five black-box queries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.