SafeSchatten: Schatten-Norm-Constrained Steepest Descent in Activation Space for Safe LLM Fine-Tuning
Abstract
Fine-tuning can improve an aligned language model's task performance while weakening its safeguards. When safety and task learning depend on overlapping representation directions, protecting safety requires controlling how those directions change during adaptation. We introduce SafeSchatten, a fine-tuning framework that budgets activation drift on trusted safety inputs and chooses updates that maximize local task progress within that budget. Frobenius and spectral constraints provide two ways to distribute the allowance across singular directions, yielding closed-form updates for the linear-layer problem. Across three downstream tasks and five aligned models, SafeSchatten substantially reduces harmfulness while retaining useful adaptation. On SST2, it reduces harmfulness from 51.30% under standard fine-tuning to 11.20%, with accuracy changing from 95.76% to 93.46%. Budget sweeps and update ablations connect these gains to the geometry of activation drift, while a one-step analysis motivates adaptive norm selection. These findings establish activation geometry as a useful design choice for safety-preserving fine-tuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.