acceptodds
Under review as a conference paper at ICLR 2027

Fine-Tuning at What Cost? Learning-Forgetting Trade-offs Across Training Algorithms

Abstract

Fine-tuning Large Language Models (LLMs) on specialized tasks often comes at the cost of degraded performance on general capabilities—a phenomenon known as catastrophic forgetting. While numerous fine-tuning methods have been proposed, systematic comparisons of their learning-forgetting trade-offs remain limited. In this work, we present the first comprehensive empirical study comparing various fine-tuning methods spanning supervised fine-tuning (SFT) variants, reinforcement learning (RL) variants, and on-policy distillation (OPD) variants. We evaluate learning gains on target tasks against forgetting measured on a diverse benchmark suite. Using Pareto frontier analysis, we characterize the achievable trade-offs for each method across learning rates and training durations. Our results reveal that no single algorithm universally dominates the learning-forgetting Pareto front, and the trend varies across models and datasets. What holds across different algorithms is that learning rates affect both learning and forgetting, while training steps primarily affect learning but not forgetting after the initial steps. Through this study, we suggest that algorithm and hyperparameter selection interact with model architecture and target dataset in ways that preclude universal recommendations, and that the learning-forgetting trade-off is better understood through Pareto front analysis than single-point comparisons.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.