Fisher-Guided Progressive Parameter Selection for Adaptive Fine-Tuning
Abstract
Parameter-efficient fine-tuning often selects trainable parameters before adaptation using architectural heuristics, without accounting for their varying importance during training. We introduce FisherAdapTune, which progressively selects parameter groups based on temporal drift in their Fisher information. Under a local Gaussian approximation, we bound the divergence between the fine-tuned posterior and pretrained prior by accumulated Fisher-weighted update costs, motivating curvature-aware selection. FisherAdapTune measures Jensen-Shannon distance between successive Fisher-value distributions and uses an adaptive threshold to freeze stabilized groups. Across VTAB-1k classification tasks, it achieves the highest macro Top-1 accuracy among the compared methods with a smaller average trainable set than full fine-tuning. Across four segmentation backbones, it maintains competitive in-distribution performance and improves zero-shot transfer in several settings. Its selections reveal architecture-dependent patterns where input, output, and normalization parameters can remain trainable as attention and MLP groups freeze at different rates. The results support Fisher structural drift as a task-dependent signal for allocating updates during adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.