Beyond Grasping: A Multimodal Human Demonstration Dataset and Benchmark for Dexterous Tool Use
Abstract
Dexterous tool use requires an embodied agent to understand task procedures, select appropriate tools and grasp types, regulate contact forces, and generate coordinated hand motions, yet these coupled capabilities remain underrepresented in existing manipulation datasets. We introduce PalmDex, a robot-free multimodal human demonstration dataset and benchmark covering 80 distinct tools and task-relevant objects, 59 real-world tasks, and approximately 31 hours of synchronized recordings, with more than 200 annotated clips per task on average. PalmDex provides synchronized egocentric and exocentric videos, 3D hand kinematics, wrist trajectories, and full-hand tactile signals. The benchmark assesses four complementary capabilities: tool-use cognition, interaction knowledge, cross-modal alignment, and motion prior learning. It comprises 43,391 question-answer pairs, together with temporal grounding, cross-modal retrieval, and motion prediction tasks. Experiments with more than 20 vision-language models and task-specific baselines show that no single model performs consistently across all four capabilities. Modality ablations reveal that vision primarily supports scene and procedure understanding, while temporally aligned tactile signals provide complementary cues for contact- and force-related reasoning. Furthermore, fine-tuning exclusively on PalmDex yields an average absolute improvement of 12.7 percentage points across four shared tasks on the independently collected EgoTouch dataset, without using any target-domain training data. These results establish PalmDex as a unified benchmark for learning and evaluating transferable interaction knowledge in dexterous tool use. The dataset, benchmark, and code will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.