FINGR: Learning Dexterous Hand Control for Real-World Rubik’s Cube Solving
Abstract
Manipulating a Rubik's Cube with one single dexterous hand is a challenging test of sustained, contact-rich control: the hand must execute successive layer turns while keeping the cube secure. Each turn requires some fingers to support the cube while others push a moving layer, release contact, and reset for the next move. To learn this coordination, we introduce (Future-supervised Interaction Network with Geometric Representations), a policy that combines finger-relative geometry with future interaction prediction. A shared point encoder expresses the cube relative to each fingertip and aggregates its points without depending on cubie indexing. Learned future tokens share the observation encoder and receive supervision for contact-force changes, layer-turn progress, and finger joint displacement at multiple time scales. The resulting representation conditions a flow policy that directly generates finger actions. On a real dexterous hand, our policy achieves around success over 300 turn attempts, compared with for the base flow policy. Integrated with grasping and table-assisted regrasping, the policy solves all ten scrambled cubes in a mean complete-system time of approximately seconds.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.