acceptodds
Under review as a conference paper at ICLR 2027

AIRP: Alignment-Induced Representation Protection against Harmful Fine-Tuning

Abstract

Fine-tuning-as-a-service adapts large language models to downstream tasks, but a small fraction of harmful examples can erode safety alignment. Current fine-tuning-stage defenses seek to preserve safety alignment through data selection, safety supervision, or constraints on model updates and representations. However, selectively applying protection remains challenging because the effects of harmful fine-tuning on safety alignment are uneven. Our analyses show that harmful examples can differ in their effects on model safety after fine-tuning, with the associated representation changes distributed unevenly across model layers. Taken together, these findings motivate Alignment-Induced Representation Protection (AIRP) for fine-grained safety preservation during downstream fine-tuning. AIRP selectively adjusts task supervision across examples and adaptively protects representations across layers, while direct refusal supervision further reinforces the model's refusal behavior. Across tasks, model architectures, harmful-data ratios, and fine-tuning set sizes, AIRP improves safety retention while maintaining competitive downstream performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.