acceptodds
Under review as a conference paper at ICLR 2027

Tangent Geometry Structures Fine-Tuning Forgetting

Abstract

Fine-tuning can degrade capabilities that a model acquired before adaptation, but it is usually unclear before training which capabilities will be affected. We study this problem in the model's local tangent space and show that forgetting has a structured geometry there. Under gradient flow, the change in old-task outputs admits an exact decomposition: the new-task loss gradient is transported into the old-task function space through the cross-task tangent kernel . This identity separates the forcing induced by the incoming task from the geometry through which it affects existing outputs. In small-output regimes, the resulting forgetting drift is strongly anisotropic and concentrates along leading eigendirections of the old-task kernel . This structure suggests a pre-adaptation diagnostic: evaluate the tangent geometry once, before any parameter update, and use it to rank which outputs or capabilities are most likely to move. We test this prediction across vision models and four language-model families. For LoRA fine-tuning of language models, the resulting unnormalised cross-task gradient score ranks held-out capability changes before training, including on three model families not used during development. In contrast, cosine-normalised gradient similarity is negatively correlated with the same realised changes across all tested configurations, indicating that the gradient magnitude removed by normalization contains information relevant to interference. The tangent-space formulation also predicts an optimizer dependence: incorporating the AdamW preconditioner improves the one-step prediction under AdamW but not under SGD, as confirmed by permutation controls. For frozen representations, we further derive a Gauss–Newton correction for cross-entropy. The curvature-aware operator predicts realised output drift better than a curvature-free predictor even when the latter is given an oracle scalar ridge parameter. When representations are allowed to change, the tangent kernel rotates and the quality of a pre-adaptation predictor degrades with training, giving it a measurable validity horizon. Finally, constraining the vulnerable eigenspace identified by the tangent geometry reduces forgetting in vision models and cross-entropy degradation in language models. Our results characterize forgetting as a structured, locally predictable phenomenon in function space. The proposed diagnostic predicts relative vulnerability rather than forgetting magnitude, and is most reliable over short adaptations for which the local tangent geometry remains informative.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.