Theory of Adaptation and Retention
Abstract
Fine-tuning (FT) should help a model learn new tasks without losing capabilities acquired during pre-training (PT). We develop a theoretical framework for understanding when these two goals can be achieved together. Using a linear prediction model, we first consider the favorable setting where both PT and FT data are available. Even in this setting, linear updates, including full FT and low-rank adaptation (LoRA), can face an unavoidable tradeoff between adaptation and retention. We then ask what the best possible correction can achieve. We characterize the minimum mean-square error (MMSE) predictor and derive upper and lower bounds on its loss that match up to a universal constant factor for arbitrary population densities. These bounds explain how population overlap and the required correction determine the minimum loss achievable by any method. We next turn to a setting closer to practice where only FT data are available. We study a practical input-dependent FT method motivated by the MMSE predictor, whose default behavior favors preserving pre-trained knowledge. With the correction held fixed, we prove that suitable separation and training conditions in a Gaussian model allow this method to achieve small FT loss while keeping forgetting small throughout training. These results explain how learning when to apply a correction can support adaptation and retention without access to PT examples.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.