Cayley-Parameterized Task-Vector Transfer: Invariant Re-basing of Fine-tuning Updates
Abstract
Task vectors capture the changes in a model's weights from fine-tuning and offer a way to reuse learned knowledge across models. Effective task-vector transfer, however, depends on how these updates are aligned and scaled to the recipient model. We introduce Cayley-Parameterized Task-Vector Transfer (CTT), which jointly learns orthogonal residual-stream rotations and scaling factors that re-base the task vectors from the original model to the recipient model. We parameterize the rotations with the Cayley transform, which keeps them orthogonal and differentiable throughout training. We evaluate CTT through three sets of experiments. First, we compare it against other state-of-the-art task-vector transfer methods on both image and text domains using CLIP-ViT B/16. On average, CTT recovers more of the domain than the other methods on the text domains, and performs comparably well to the other methods on the image domains. Then, we apply CTT to natural language autoencoders (NLAs), which describe model activations in natural language and reconstruct them from those descriptions. Since NLAs degrade when their underlying base models are fine-tuned and struggle to reconstruct out-of-distribution activations, task-vector transfer provides an effective way to restore them without full-parameter retraining. We find that CTT consistently recovers more of the NLA domains than the other methods. Finally, we compare CTT directly with GRPO and SFT fine-tuning of the NLA. Under the same training-data distribution, CTT achieves comparable code-domain recovery and higher support retention while using significantly fewer estimated training FLOPs than GRPO, and achieves greater code-domain recovery than SFT. Together, these results show that CTT is the only method among those we compare that respects reparameterization symmetries and transfers task vectors effectively across models. The differences that we observe between the image and text domains also suggest that the effectiveness of task-vector transfer methods may be modality-specific.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.