Rolebind: training vision-language-action policies with task-conditioned role–instance supervision
Abstract
Vision-Language-Action (VLA) models have demonstrated strong capabilities in general-purpose robotic manipulation by jointly modeling visual observations, language instructions, and robotic actions. However, existing VLA models primarily rely on action supervision for policy learning, and there is a lack of explicit constraints on the correspondence between task semantics and physical entities. Robotic manipulation requires not only the identification of relevant entities but also an understanding of their roles in the current task. To address this issue, we propose RoleBind—a training method for implementing the binding of task roles to object instances within existing VLA policies. In addition to the original action learning objective, RoleBind introduces role-instance binding supervision, which explicitly aligns task roles (such as manipulated objects and target objects) with their corresponding visual instances in a shared representation space, thereby enhancing the VLA’s representation of task-relevant entities and their role relationships. Role labels and instance region information are used only during the training phase. We applied RoleBind to different action modeling methods on LIBERO and further evaluated its generalization ability in out-of-distribution scenarios. Experimental results show that RoleBind increases the average success rate of StarVLA-OFT from 98.0% to 98.55%, an improvement of 0.55 percentage points, and yields a further performance improvement of 0.8 percentage points on StarVLA-. These results demonstrate that RoleBind can improve task success rates without altering the original inference process, and also provide a viable approach for explicitly linking task semantics to physical entities in VLA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.