acceptodds
Under review as a conference paper at ICLR 2027

CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models

Abstract

This paper proposes a novel approach to address the challenge that pretrained vision-language-action models (VLAs) often fail to effectively improve performance and reduce adaptation costs during standard supervised finetuning (SFT). Some advanced finetuning methods with auxiliary training objectives can improve performance and reduce the number of convergence steps. However, they typically incur significant computational overhead due to the additional losses from auxiliary objectives. To simultaneously achieve the enhanced capabilities of auxiliary training with the simplicity of standard SFT, we propose a new paradigm for steering the behavior of VLAs, centered around capability vectors. Capability vectors specify the beneficial properties induced by auxiliary-objective SFT and are given by element-wise arithmetic operations between the auxiliary-objective SFT models and regular SFT models. The capability vectors are then merged with pretrained parameters to form a capability-enhanced meta model. With a regular SFT integrated with a lightweight orthogonal regularization loss, the capability-enhanced meta model attains performance comparable to auxiliary finetuned baselines with reduced computational overhead and higher training efficiency. Internal and external experiments demonstrate that our capability vectors (1) are effective and versatile across diverse models, (2) can generalize to novel environments and embodiments out of the box.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.