TCR-ARM: Tactile Context Reasoning and Action Refinement for Contact-Rich Robotic Manipulation
Abstract
Contact-rich robotic manipulation depends on touch to resolve contact states that vision alone may fail to reveal. Yet tactile information should not influence control uniformly. Persistent tactile context can support reasoning about contact state and interaction progress, whereas tactile-driven local action refinement should be applied selectively. We present TCR-ARM, a vision-language-action policy that realizes this distinction through two tactile interfaces derived from a shared encoder. Compact tactile representations support persistent tactile context reasoning in the vision-language model, while spatially dense tactile features adaptively update wrist visual tokens through a learned gate to support local action refinement. A learnable gate alone, however, does not ensure selective use of dense refinement. We therefore introduce route-aware tactile pretraining (RAT), which samples no-tactile, context-only, and full-tactile routes to train tactile context reasoning with and without tactile-driven action refinement. After pretraining on 1,500 hours of tactile manipulation data and post-training with 100 demonstrations per task, TCR-ARM achieves 70.0% average success across six real-world contact-rich robotic manipulation tasks, outperforming the strongest evaluated baseline by 15.0 percentage points. RAT further improves average success by 10.8 percentage points over full-tactile-only training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.