AIRP: Alignment-Induced Representation Protection against Harmful Fine-Tuning
Abstract
Fine-tuning-as-a-service adapts large language models to downstream tasks, but a small fraction of harmful examples can erode safety alignment. Current fine-tuning-stage defenses seek to preserve safety alignment through data selection, safety supervision, or constraints on model updates and representations. However, selectively applying protection remains challenging because the effects of harmful fine-tuning on safety alignment are uneven. Our analyses show that harmful examples can differ in their effects on model safety after fine-tuning, with the associated representation changes distributed unevenly across model layers. Taken together, these findings motivate Alignment-Induced Representation Protection (AIRP) for fine-grained safety preservation during downstream fine-tuning. AIRP selectively adjusts task supervision across examples and adaptively protects representations across layers, while direct refusal supervision further reinforces the model's refusal behavior. Across tasks, model architectures, harmful-data ratios, and fine-tuning set sizes, AIRP improves safety retention while maintaining competitive downstream performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.