Reparameterize Before You Extract: Training-Free Spectral Preconditioning for Knowledge Transfer
Abstract
Large-scale pre-trained models contain transferable knowledge, but their parameters are tied to the architecture and size of the source model, limiting direct reuse across model sizes. Existing learngene-based approaches seek compact, task-agnostic representations for variable-sized model initialization, but many require auxiliary training. FRONT shows that low-frequency components of pre-trained weights are informative for foundational and transferable knowledge. However, standard pre-training does not explicitly organize weights for subsequent knowledge extraction, and the resulting spectra can exhibit limited low-frequency concentration. Under a fixed coefficient budget, this constrains the amount of spectral energy that can be retained. We introduce Spatial Permutation Alignment (SPA), a training-free structural preconditioner for frequency-domain knowledge extraction. Our key observation is that functionally equivalent parameterizations within the same permutation-equivalence class can exhibit different spectral energy distributions. SPA exploits this reparameterization freedom through depth-wise bipartite matching and width-wise anchored joint permutation, reducing inter- and intra-layer structural variation. Integrating SPA with FRONT yields SPA-FRONT, which captures a larger fraction of spectral energy in the retained low-frequency coefficients under the same coefficient budget. Across vision and language models, SPA-FRONT substantially improves variable-sized model initialization, reducing vision-model training epochs to reach target accuracy by 30.0% on average and cutting language-model pre-training FLOPs by up to 58.7%. The code is available at https://anonymous.4open.science/r/SPA-FRONT.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.