Exec-VTLA: Execution-Aware Vision-Tactile-Language-Action with Measured Motion and Contact Histories
Abstract
Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but the effects of an action occur at the contact interface, which often lies in the blind spot of visual observation. Without sufficient information about actual motion and contact changes, policies may continue motions that are no longer appropriate after contact, leading to manipulation failures. To address this limitation, we present Exec-VTLA, an execution-aware vision-tactile-language-action policy that conditions future actions on measured motion and contact histories. Its core is a a multi-source action head that uses future action queries. Each future action query jointly reads four complementary conditions through shared attention: visual-language context for task semantics, robot state for the current configuration, measured state transitions for the motion that actually occurred, and bilateral temporal tactile observations for changes at the contact interface. The action head jointly predicts an action chunk relative to a common reference state, incorporating recent execution information into the generation of the next action segment. On the UniVTAC benchmark, Exec-VTLA achieves an average success rate of 68.17% across six contact-rich tasks. It further achieves an average success rate of 80.00% across four tasks on a physical robot platform. Ablation studies support the effectiveness of measured motion history, tactile input, and the proposed multi-source fusion. These results demonstrate the effectiveness of integrating motion and contact histories into action decoding for contact-rich manipulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.