Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning
Abstract
Released aligned large language models remain vulnerable to malicious downstream fine-tuning. Existing defenses are largely designed for the fine-tuning-as-a-service (FTaaS) paradigm or rely on downstream users to follow additional safety procedures, and therefore do not directly address the setting we study, where a provider releases the model for downstream fine-tuning. We adopt a partially protected open-weight (PPOW) release, in which most weights remain open and trainable while a small safety-critical component is preserved at release. We propose a Unidirectional Safety Gate (USG), instantiated as a Null Space Cubic Layer together with an Inverse Adapter inserted after the final Transformer layer. During downstream fine-tuning, the cubic layer suppresses or blocks gradients from harmful samples whose hidden states fall in a calibrated protected region, while the Inverse Adapter restores the base model's forward behavior. In practice, the release threshold is calibrated once on defender-held harmful reference data, and blocking generalizes beyond the build set for coherent harmful distributions. At this defender-controlled in-coverage operating point, USG holds post-finetuning attack success rate within 0.02 of its pre-release level across six model–dataset settings, while passing benign data and leaving continued safe fine-tuning intact. Because the cubic forward is analytically invertible, and hence removable from a purely open-weight release, we target removal-resistance as a designed property through a root-of-trust protected-release structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.