acceptodds
Under review as a conference paper at ICLR 2027

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

Abstract

Released aligned large language models remain vulnerable to malicious downstream fine-tuning. Existing defenses are largely designed for the fine-tuning-as-a-service (FTaaS) paradigm or rely on downstream users to follow additional safety procedures, and therefore do not directly address the setting we study, where a provider releases the model for downstream fine-tuning. We adopt a partially protected open-weight (PPOW) release, in which most weights remain open and trainable while a small safety-critical component is preserved at release. We propose a Unidirectional Safety Gate (USG), instantiated as a Null Space Cubic Layer together with an Inverse Adapter inserted after the final Transformer layer. During downstream fine-tuning, the cubic layer suppresses or blocks gradients from harmful samples whose hidden states fall in a calibrated protected region, while the Inverse Adapter restores the base model's forward behavior. In practice, the release threshold is calibrated once on defender-held harmful reference data, and blocking generalizes beyond the build set for coherent harmful distributions. At this defender-controlled in-coverage operating point, USG holds post-finetuning attack success rate within 0.02 of its pre-release level across six model–dataset settings, while passing benign data and leaving continued safe fine-tuning intact. Because the cubic forward is analytically invertible, and hence removable from a purely open-weight release, we target removal-resistance as a designed property through a root-of-trust protected-release structure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.