acceptodds
Under review as a conference paper at ICLR 2027

Learning-Forgetting Optimality in Supervised Finetuning: A Cliff Perspective

Abstract

Supervised finetuning (SFT) of pretrained language models trades off the acquisition of new domain capabilities against retention of prior knowledge. In this work, we extensively characterize this Pareto frontier of learning versus forgetting across nine instruction-tuned models from the Qwen2.5, Llama-3, and OLMo-2 families with full learning-rate sweeps for a number of classical and modern methods for continual learning and domain adaptation. Doing so, we find that across models, this frontier almost universally develops into a sharp cliff: a sharp phase transition where training recipes flip from learning without forgetting into catastrophic forgetting without improvements in learning. Across all methods we test, EWC, MAS, MIGU, NanoAdam, KL regularization, LoRA, and two geometry-motivated baselines we propose, we find that, while many methods can improve retention if the SFT baseline does not yet optimally retain, improvements in learnability remain elusive in practice for modern LLMs. This delineates two regimes. Either baseline SFT performance appears as a gradual trade-off between learning and forgetting, in which case many methods can be applied to ”sharpen” the cliff and reduce forgetting, or, baseline SFT is already not forgetting, in which case no method we test can substantially intervene.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.