acceptodds
Under review as a conference paper at ICLR 2027

Fine-Tuning Beyond Weight-Space Geometry: One Frontier, Two Origins

Abstract

Post-training methods are commonly evaluated through the tradeoff between new-task performance and retention, yet the shape of this frontier does not reveal its cause. A similar loss of adaptation can be either necessary, because the old and new tasks demand incompatible predictions, or avoidable, because an old-data-free objective constrains the wrong computations. We expose this ambiguity using controlled causal sequence tasks with known target distributions and replay solutions. We study WindowLoRA, a family that moves from ordinary LoRA to dense teacher distillation by progressively increasing the causal and positional coverage of hidden-state matching. Across both conflicting and compatible tasks, increasing constraint density produces the same qualitative continuum: old-task error decreases while new-task error increases. The meaning of this continuum, however, changes relative to replay. Under shared-input target conflict, dense output distillation reaches the replay compromise, showing that the remaining tradeoff is intrinsic. Under compatible disjoint support, replay solves both tasks while new-input matching remains on a nonzero frontier, showing that the tradeoff is introduced by incomplete preservation supervision. We then test this interpretation in pretrained decoder-only models by comparing the representation drift and downstream retention–accuracy frontiers induced by different adaptation methods. Although stronger constraints reduce drift, none eliminates the frontier, and WindowLoRA provides controlled intermediate operating points. Finally we position the family to K-prior which explains how a new-input matching can replace the unavailable old objective only when its probes cover the old-sensitive directions of the adaptation carrier.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.