Seeing Heat, Acting by Touch: Thermally Grounded Prompting for Vision-Tactile-Language-Action Models
Abstract
Vision-Tactile-Language-Action (VTLA) models augment Vision-Language-Action (VLA) policies with tactile feedback, enabling robots to respond to contact deformation, slip, and grasp stability during physical interaction. However, tactile sensing is available only after contact and cannot safely identify pre-contact target properties such as object temperature or electrical energization. A direct solution is to inject both thermal and tactile tokens into a pretrained action backbone. With limited robot demonstrations, however, jointly learning heterogeneous representations, cross-modal alignment, and continuous control from action supervision is data-inefficient and prone to RGB-dominated modality imbalance, potentially suppressing thermal and contact-dependent tactile cues. Scaling synchronized RGB–thermal–tactile–language–action data to alleviate this problem is also substantially more resource-intensive than collecting conventional RGB-based robot demonstrations. To address these challenges, we propose Thermo-VTLA, a hierarchical framework that separates thermal physical-state grounding from contact-rich VTLA control. A thermally grounded vision-language model receives paired RGB–thermal observations and the original instruction, and produces a constrained executable prompt. A prompt-conditioned VTLA controller then executes the command using RGB observations, robot state, and visuotactile feedback, without receiving thermal tokens. To effectively incorporate tactile feedback without disrupting the pretrained vision-language representation, we introduce a two-stage contact-aware tactile integration strategy. Contact-aware gating first determines tactile validity from the interaction state, after which gated fusion adaptively regulates the contribution of valid tactile representations to the action context. This decomposition preserves tactile sensing in the high-frequency control loop while moving thermal reasoning to an explicit semantic interface. We evaluate Thermo-VTLA on beverage retrieval, bread manipulation, and electrical-plug disconnection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.