CORE: Identifying and Protecting Critical Outliers in LLMs against Functionality Stealing
Abstract
Protecting proprietary large language models (LLMs) on user-controlled devices requires preventing unauthorized reuse while preserving normal inference. We investigate which components to prioritize based on their importance to model functionality and how to hinder their reconstruction from observable input–output pairs. We find that removing a small subset of outlier-associated weights from the MLP of one specific decoder layer severely degrades model performance, whereas removing the same number from other layers has much less effect. This finding motivates prioritizing the corresponding critical MLP for protection. We propose CORE, a learning-based defense algorithm that combines fixed encodings with learned selective recovery to hinder reconstruction of this component. With the original LLM parameters frozen, the recovery module is trained to support normal inference, while occasional output modifications shift the MSE-optimal surrogate prediction away from the output used in most executions. The fixed encodings and trained module can be reused without preparing fresh masks and corresponding recovery values for each inference forward pass. Across three LLMs, we show that the proposed CORE retains 96.0-98.8% of the original average zero-shot accuracy, while an MSE-trained surrogate replacement performs 17.8-31.2% below the protected model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.