Not All Forgetting Is Equal: Budgeted Preference Optimization for Retaining Critical Capabilities
Abstract
Preference optimization is known to erode abilities acquired earlier, and existing remedies treat this alignment tax as a single quantity to bound. Tracking six capabilities through full-parameter DPO of Qwen2.5-1.5B-Instruct over three seeds, we find instead that the tax is paid unevenly. Held-out safety falls by up to 5.3 points, whereas instruction following and truthfulness improve and the remaining capabilities hold. As a result, an average over capabilities rises and hides the safety loss. Reference-free SimPO erodes math as well. Moreover, likelihood margins between the model’s own verified and failed safety answers keep rising, so the most readily available retention signal misses the loss. Motivated by these findings, we propose RETAINPO, which separates detecting capability loss from protecting against it. RETAINPO measures a statistically supported loss for each capability on task-native sentinel questions, spends a capped token budget on the capabilities with the largest supported loss and abstains when none is supported. A hinge on the likelihood margins of verified certificates from those capabilities then enters the preference objective. Across three seeds, RETAINPO restores held-out safety from 92.5 to 97.0 percent at a preference accuracy comparable to DPO and keeps its truthfulness gain, whereas every anchoring baseline that retains safety gives that gain up. Under SimPO, it adds math once math loss appears and reaches a smaller worst-capability loss than the tested fixed allocations, although it spends more retention tokens. A protection-transfer matrix further shows that protecting safety raises its held-out safe rate by 5.9 points over DPO and exceeds the average effect of protecting the other five capabilities by 5.0 points. Certificates thus fail as a detector yet work as an intervention, and our findings recast retention during preference optimization as an online allocation problem that intervenes only on measured capability loss.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.