acceptodds
Under review as a conference paper at ICLR 2027

SafeSchatten: Schatten-Norm-Constrained Steepest Descent in Activation Space for Safe LLM Fine-Tuning

Abstract

Fine-tuning can improve an aligned language model's task performance while weakening its safeguards. When safety and task learning depend on overlapping representation directions, protecting safety requires controlling how those directions change during adaptation. We introduce SafeSchatten, a fine-tuning framework that budgets activation drift on trusted safety inputs and chooses updates that maximize local task progress within that budget. Frobenius and spectral constraints provide two ways to distribute the allowance across singular directions, yielding closed-form updates for the linear-layer problem. Across three downstream tasks and five aligned models, SafeSchatten substantially reduces harmfulness while retaining useful adaptation. On SST2, it reduces harmfulness from 51.30% under standard fine-tuning to 11.20%, with accuracy changing from 95.76% to 93.46%. Budget sweeps and update ablations connect these gains to the geometry of activation drift, while a one-step analysis motivates adaptive norm selection. These findings establish activation geometry as a useful design choice for safety-preserving fine-tuning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.