Optimal Transport Layer Alignment for Cross-Architecture Task Vector Re-basin
Abstract
Task vectors encode task-specific adaptations, but transferring them across architectures is difficult when source and target models differ in width and depth. Existing re-basin methods address width mismatches through coordinate alignment, yet handle depth mismatches using fixed layer correspondences. We formulate task vector re-basin using correspondence weights to combine source representations for alignment and task vector blocks for transport. To estimate these weights, we introduce Optimal Transport Layer Alignment (OTLA), a training-free method that learns adaptive, weighted correspondences between layers of heterogeneous models. Using a small unlabeled calibration set, OTLA combines representational similarity measured by Centered Kernel Alignment with a relative-depth prior in a relaxed optimal transport formulation. The resulting transport plan jointly determines which source layers contribute to each target layer. The weighted formulation supports expansion and compression without modifying the source architecture or fine-tuning the target model. Across eight vision tasks, three Vision Transformer scales, and both model expansion and compression, OTLA achieves the highest average accuracy when integrated into existing re-basin methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.