What Makes a Layer Adaptable? A Mechanistic Study of Language Model Fine-Tuning
Abstract
What makes one Transformer layer more useful to fine-tune than another? We judge a layer by the loss-reducing output change, or correction, it can produce. Our first characterization separates the correction a layer can express from the spectral strength that determines how readily it is learned. These factors can trend in opposite directions with depth and produce a nearly flat gradient profile, so gradient size can hide clear differences in what layers can correct. Our second characterization shows how joint training redistributes the remaining error and available output directions across layers; as a result, when earlier layers are frozen, the retained layers can absorb part of the correction the frozen layers would otherwise have made. Controlled GPT-2 experiments test these mechanisms, and a systems study establishes the cost advantage of placing trainable layers nearest the output. Across 15 datasets on each of Qwen3-8B and LLaMA2-7B, most new loss-relevant output change occurs early in training, and the share of accumulated change covered by each output-nearest layer set remains stable over training within a dataset but differs strongly across datasets. These results yield a roadmap parameterized by when to stop broad training and how many output-nearest layers to retain, spanning all-layer, focus-only, and broad-to-focused tuning under different resource budgets. One implementation, Suffix95, trains across all layers until a chosen step and then retains the smallest output-nearest layer set covering % of the measured all-layer change. Across 39 model and dataset combinations, it achieves competitive median performance relative to full-depth baselines while reducing total training compute by –% and post-transition memory by –%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.