What Makes an Update Harmful? The Dynamics of Functional Risk in Model Adaptation
Abstract
What makes an adaptation update harmful to what a model already knows? Existing preservation methods constrain updates using historical gradients, activation-derived directions, or protected subspaces. We take a more direct functional view: for a protected loss, we ask how a concrete update proposed by the optimizer changes it. To first order, this change is the inner product between the update and the gradient of the protected loss, which we call the functional risk of the update. Estimating this gradient from finite reference data raises two questions for any method built on such estimates: how accurately must the signal be estimated to judge an update, and how long does it remain informative as the model changes? Judging an update requires only the gradient component along it, but AdamW proposals are nearly orthogonal to the protected gradient, and we prove that any procedure then needs a number of reference examples that grows with the inverse squared cosine between the two. Correcting an update is easier: a correction along the estimated gradient lowers the protected loss whenever the estimate is aligned with the true gradient. Across model states, a cached estimate changes later decisions, most strongly after a task switch. Motivated by these findings, we introduce Dynamic Functional Risk Control (DFRC), a lightweight controller that periodically refreshes its estimates and minimally modifies candidate updates whose predicted risk exceeds a tolerance; its forgetting over a run is bounded independently of detection accuracy. In sequential adaptation of two model families over three task orders, DFRC improves average accuracy and reduces forgetting relative to sequential LoRA, and further reduces forgetting when combined with an existing LoRA continual-learning method.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.