Resistance Training to Reduce Forgetting
Abstract
Task adaptation can improve a language model's target-task performance while weakening other important behaviors, even under on-policy training. We introduce Resistance Training (ResT), which reduces this forgetting through a temporary behavioral load. We fit a low-rank adapter to an unwanted behavior with the backbone frozen, then freeze the adapter while the backbone learns the task. Removing the adapter leaves a model with no additional inference cost. Across medical and financial adaptation of Qwen3-4B, targeted loads improve safety refusal or factual restraint over vanilla on-policy distillation and self-distillation while largely retaining task gains. Under MedQA on-policy distillation, ResT lowers the mean hallucination score by 40 points (0–100 scale), at comparable task accuracy. How the load is fitted also matters: on-policy distillation gives stronger hallucination protection than response supervision under both downstream distillation and supervised fine-tuning. This identifies resistance construction as a design choice for behavioral retention beyond the downstream objective. A linearized analysis gives conditions under which task learning compensates for the load, improving the protected behavior after removal.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.