acceptodds
Under review as a conference paper at ICLR 2027

What Should We Freeze? Guarded Freezing: Connectivity Shapes the Fine-Tuning of Pretrained Models

Abstract

When adapting pre-trained models through fine-tuning, freezing weights alone might not preserve performance, as updates elsewhere can change the inputs to the frozen core, ultimately affecting overall performance. We first analyse the case where a selected frozen core can be isolated and propose _removal-value_, a capacity-based score that approximates HOPE's removal cost averaged over removal orders. We show that in VGG-8, cutting paths from trainable neurons into a frozen core makes selection using this score useful: frozen preserves percentage points more old-task accuracy than DEFT at similar new-task accuracy. In transformers, shared residual streams leave paths into frozen neurons open. For this case, we derive _drift-value_, a forward-only proxy for the output disturbance from updating each weight entry under a local update model. In language models, at 40 epochs, this policy exceeds adapted Wanda and RIA freezing scores in settings with substantial retention loss, while its differences from Fisher remain unresolved. After 160 epochs on Qwen2.5-1.5B, it retains more than static Fisher. In DINOv3 vision-transformer adaptation to point clouds, drift-value retains image accuracy versus for a random mask of the same count. These results motivate _Guarded Freezing_: select by removal-value when incoming paths are cut, and by drift-value when they remain.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.