ReDexT: Learning a Generalizable Residual Policy for Dexterous Retargeting
Abstract
Dexterous retargeting often requires fitting robot controls to each human hand-object motion, making large collections expensive to process. Offline inverse kinematics provides inexpensive nominal controls, but cannot correct errors caused by contact and object dynamics. We present ReDexT, a two-stage reinforcement-learning framework that learns one residual policy to execute new trajectories from inverse-kinematics commands. The key is to learn reusable feedback by varying both the commands the policy corrects and the motions it practices. We first perturb successful source commands while keeping their motion targets fixed, so the policy learns to correct varied execution errors. Rollout-guided training then broadens trajectory coverage through fresh on-policy learning, without treating rollout actions as imitation labels. Once trained, the shared policy executes new trajectories with frozen weights and supports optional local or shared adaptation for better tracking. On new trajectories screened for initial stability in single-hand simulation, frozen ReDexT achieves higher success and shorter processing times than the evaluated per-trajectory methods. Shared adaptation further raises success from 55.08% to 80.08% across 256 test trajectories. Ablations show that control perturbations ease transfer to kinematic commands, and expanded IK training broadens success coverage. The training recipe extends to four additional robot hands, each with a separately trained policy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.