DexWorld-Bench: Evaluating Fine-Grained Control and Interaction for Dexterous World Models
Abstract
World models for dexterous manipulation have served as learned simulators for anticipating action consequences and supporting robotic planning and policy learning. Their utility depends on faithfully capturing fine-grained finger motions and hand-object interactions, where small control differences can change manipulation outcomes. However, existing evaluation methods provide limited systematic coverage of these dynamics. Similar overall motion or successful task completion can conceal errors in local control and interaction. We introduce DexWorld-Bench, a benchmark that evaluates dexterous world models along three complementary dimensions: fine-grained control accuracy, interaction consistency, and task executability. The benchmark comprises 220 full task sequences, 90 critical interaction clips that include both successful and failed interactions, and 70 controlled action comparisons that isolate the motion of a single arm or finger. We evaluate multiple widely-used world models spanning general-purpose video generators, action-conditioned world models, and world action models under the same protocol. The results show that action conditioning alone does not ensure that models faithfully follow fine-grained action commands, and that fine-grained control and long-horizon executability remain key bottlenecks in current dexterous world models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.