Tac-DINO: Learning Tactile Features with Patch Alignment
Abstract
Multisensory perception is essential for robotic manipulation, with tactile sensing providing critical information for fine-grained, contact-rich interactions. Although prior work integrates tactile and spatial signals to improve perception, vision-tactile learning commonly relies on object-level alignment. Addressing this limitation requires systematically aligned 3D vision-tactile data. We develop a data collection system and introduce Touch3D, a dataset comprising 505 real-world objects and 20,025 tactile contacts. Building on this dataset, we establish the Contact-Centric Multimodal Tactile Representation Benchmark (CoMT-Bench), which expands evaluation from object-level retrieval and tactile properties to include patch-level contact evaluation. We further propose Geometry-Relation Distilled Tactile-Centric Cross-Modal Learning (GeoTAC), which explores tactile learning across all multisensory data. Experiments on the Touch3D test set, patch-level alignment improves the average score by 5.32 over object-level alignment, with GeoTAC's geometric supervision improving 4.48 more. In robotic manipulation, GeoTAC improves TP by 19.14 and 11.34 over DINOv2 (Vision-only) and DINOv2 (Vision-Tactile), respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.