Grounding Vision, Touch, and Sound for Contact-rich Dexterous Manipulation
Abstract
Vision, touch, and sound provide complementary cues for contact-rich dexterous manipulation, yet their distinct spatial and temporal structures make effective fusion challenging. Moreover, requiring every demonstration to contain all modalities limits the amount of usable training data. We introduce DexMuse, a plug-in multisensory fusion module that anchors touch, sound, and local vision at the fingertips of a dexterous hand, where contact most often occurs. Calibrated fingertip projections associate local visual features with tactile and proprioceptive context, while the resulting region tokens attend to acoustic history and complement global visual observations. We further propose availability-aware mixed training, which lets demonstrations missing touch or sound supervise the same policy without placeholder tokens or synthesized signals. We integrate DexMuse into both ACT and the pretrained vision-language-action model π0.5 through backbone-specific interfaces. Across seven challenging real-world single-arm and bimanual tasks, DexMuse reaches average success rates of 0.70 with ACT and 0.80 with π0.5, compared with 0.35 and 0.52 for concatenation-based fusion, and outperforms all baselines on every task. At a fixed budget of complete demonstrations, adding incomplete demonstrations improves DexMuse by 0.19 on average, more than twice the gain of concatenation-based fusion. Real-world videos are available at dexmuse-submission.github.io.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.