acceptodds
Under review as a conference paper at ICLR 2027

Do Pretrained Spectral Subspaces Matter For Adaptation?

Abstract

Do accuracy gains require updates along principal pretrained directions, and what does interaction between spectral regions change? We use fixed-band adapters to separate spectral location from within-band flexibility, then study leading–tail interaction during training and at trained checkpoints. Under fixed recipes, leading, middle, tail and random supports all improve accuracy on RTE and CommonsenseQA. These supports are not interchangeable: tail gives the lowest decoder negative log-likelihood (NLL) within each adapter family. On encoder tasks, an interaction penalty sharply reduces cross-band energy share relative to a roughly norm-matched control. It lowers uncalibrated NLL without a resolved accuracy gain; temperature scaling removes most of the NLL advantage. In a separate decoder study, removing learned cross-region components raises choice NLL relative to shrinking each matrix's update to the same remaining norm. Thus, accuracy gains do not require principal support, while the predictive effects of cross-region interaction depend on the metric and the intervention.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.