T-Rex-RL: Improving Tactile-Reactive Dexterous Manipulation with Real-World Reinforcement Learning
Abstract
Tactile-reactive vision-language-action (VLA) models combine fast closed-loop tactile control with the broad visuomotor priors acquired through large-scale pretraining. However, existing models are mainly trained with imitation learning on teleoperated demonstrations, limiting their ability to improve beyond the quality and coverage of offline data. In this work, we study how tactile-reactive VLAs can improve via real-world interaction using offline-to-online reinforcement learning (RL). To achieve this, we initialize a tactile-reactive actor and critic through offline RL with teleoperation data, then continuously update them using online robot rollouts. The VLA is pretrained on large-scale data, whereas RL post-trains it on limited task-specific robot experience. We therefore freeze the VLA and add the RL actor and critic as Mixture-of-Transformers (MoT) experts that attend to its representations. While the VLA predicts action plans at a low frequency, the RL actor refines them at a high frequency from tactile signals to react to contact changes. The critic evaluates the refined actions to guide the actor's learning. We evaluate our approach on a suite of long-horizon bimanual dexterous manipulation tasks, where our method substantially outperforms baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.