acceptodds
Under review as a conference paper at ICLR 2027

A Unified NTK-based Framework for Forgetting Approximation in Full Fine-Tuning LLMs

Abstract

Full fine-tuning enables large language models (LLMs) to acquire new knowledge but also alters previously mastered behavior. Therefore, understanding how forgetting arises during fine-tuning is crucial. We develop a neural tangent kernel (NTK)-based framework to locally approximate output changes associated with forgetting. The derivation separates the change into Gauss–Newton (GGN) quadratic term and linearization remainder terms. We use GGN quadratic to approximate the output change, which consists of three factors: transfer geometry, curvature response, and residual signal. We further show that three major forgetting-mitigation approaches can be viewed as modifications to these factors. Based on this framework, we proposed NTK-Replay, which selects replay strength by balancing approximated drift against replay penalty. Extensive experiments covering three modalities, five datasets, and twelve models validate the effectiveness of our NTK-based approximation framework.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.