Align What Forms the Update: Bilinear Factor Alignment for Cross-Model Task-Vector Transfer
Abstract
Adapting large-scale pre-trained models to specialized tasks has been one of the prevalent paradigms in modern machine learning, with fine-tuning being one of the most widely studied approaches. However, when a new version of a pre-trained model becomes available, expertise acquired through fine-tuning cannot be directly reused because it is tied to the parameterization of the original model, requiring another costly fine-tuning. To address this inefficiency, recent work seeks to transfer expertise across models in the form of task vectors—defined as the parameter difference between a fine-tuned model and its base model. However, the connection between task-vector formation and cross-model alignment remains underexplored. In this work, we revisit task-vector formation and derive an alignment principle directly from the structure of the parameter update. Specifically, each layer-wise parameter gradient factorizes as an outer product between an input-side activation and an output-side gradient, revealing the bilinear structure underlying the update. This structure exposes two distinct factor spaces and motivates aligning their coordinate systems across models. Building on our observation, we propose BIFA (Bilinear Factor Alignment), a training-free framework that separately aligns the activation and gradient factor spaces across source and target models to transfer task vectors. Across extensive vision and language-model benchmarks, BIFA consistently outperforms existing transfer methods across models that differ in width, depth, and pre-training configuration, further extending to cross-family LLM transfer with greater architectural mismatch.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.