ASCENT: First-Order Optimal Fine-Tuning with Recalibration for Safety–Utility Co-Enhancement
Abstract
Supervised fine-tuning can substantially improve the downstream utility of large language models (LLMs) but can also compromise their safety capabilities. Recent safety-preserving fine-tuning methods aim to prevent safety degradation by identifying safety-related parameters or subspaces and constraining downstream updates accordingly. However, existing methods still have several limitations. First, these methods mainly focus on safety preservation, rather than seeking joint safety and utility enhancement. Second, there is no theoretical characterization of the optimal safety-related subspace or the safety-preserving task update. Third, existing methods typically rely on a static safety subspace throughout downstream fine-tuning, even though this subspace may be outdated as the LLM evolves. To address these limitations, we propose ASCENT, a downstream fine-tuning framework for safety–utility co-enhancement through first-order optimal safety-aware periodic calibration and task optimization. Specifically, we model safety as a function of LLM parameters and use its first-order approximation to characterize how safety changes under parameter updates. Under a fixed rank and Frobenius-norm budget, we prove that the parameter update constructed from the top- singular components of the safety-function gradient maximizes the estimated safety change, and use this update as a periodic calibration to preserve and even improve the safety capability. We further formulate an optimization problem for the safety-preserving task update that minimizes its deviation from the original task update while penalizing its negative effect on the estimated safety change, which yields a unique solution for the task optimization process. As the optimization proceeds, ASCENT iteratively applies the optimal safety-preserving task update and the optimal safety-improving calibration update, thereby enabling the co-enhancement of safety and utility. Extensive experiments across multiple LLM families and downstream tasks demonstrate that ASCENT can improve downstream utility by up to 20.3% and reduce attack success rate by up to 35.5%. Compared with existing safety-preserving fine-tuning methods, ASCENT achieves state-of-the-art safety and utility across all evaluated settings. Our code is available at https://anonymous.4open.science/r/ASCENT-Code-8E96/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.