Rethinking Anti-distillation of Large Language Models: An Unlearnable Examples Perspective
Abstract
Large language models (LLMs) expose valuable capabilities through their generated responses, but making them vulnerable to unauthorized distillation. Existing anti-distillation methods largely assume that knowledge transfer relies on fitting to teacher-generated responses and therefore seek to make those responses difficult to learn. We revisit this premise and ask: *Can a student easily fit protected responses yet fail to acquire the generalizable capability behind them?* Inspired by unlearnable examples (UEs), we decouple data fitting from knowledge transfer by injecting easy-to-fit perturbations into the teacher's reasoning, encouraging the student to learn the surface shortcut while neglecting the underlying capability. Building on this principle, we propose DistilShield, which protects LLMs across the distillation trajectory: it leverages an *Unlearnable Converter* that implants implicit shortcuts into teacher reasoning, steers the learning dynamics away from capability transfer during optimization, and consolidates the induced behavior under the student's own rollout distribution. Extensive experiments demonstrate that DistilShield provides a plug-and-play defense layer that effectively suppresses knowledge distillation while preserving the utility of teacher responses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.