Shallow-Wide Fine-Tuning for Continual Learning with Pre-Trained Models
Abstract
Full fine-tuning of Pre-Trained Models (PTMs) on continually arriving downstream tasks often leads to catastrophic forgetting. Existing PTM-based Continual Learning (CL) methods typically freeze the backbone and attach auxiliary modules to preserve previously learned knowledge. However, frozen representations limit plasticity under large distribution shifts, and auxiliary modules incur additional inference overhead. To address these limitations, we propose Shallow-Wide Fine-Tuning (**SW-Tune**), which enables stable backbone updates without any auxiliary components. Inspired by the empirical insight that shallow-and-wide architectures exhibit inherently greater resilience to forgetting than deep-and-narrow ones, SW-Tune introduces a principled optimization framework that modulates the gradient subspace of a deep PTM to emulate shallow-and-wide learning dynamics without physical architectural changes. Specifically, *periodic block freezing* restricts trainable parameters to an equidistant, sparse subset of blocks, functionally reducing the effective trainable depth. Simultaneously, *gradient spectral flattening* regularizes the singular-value spectrum of layer-wise gradients, compelling parameter updates to span a broader orthogonal subspace, thereby expanding the effective capacity width. Extensive experiments demonstrate that SW-Tune outperforms existing PTM-based CL methods while introducing no additional inference overhead. Code is available at https://anonymous.4open.science/r/SW-Tune
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.