SetMimic: Policy-in-the-Loop Interpretation of Motion Intent for Humanoid Imitation from Imperfect Motion References
Abstract
Humanoid imitation commonly relies on high-quality motion capture of professional performers, treating reference joint trajectories as reliable targets for tracking rewards. Internet videos offer a much broader source of motion data, but video-based motion capture often produces physical artifacts such as foot skating, ground penetration, and inconsistent contacts. Directly rewarding joint-level tracking of these references can encourage policies to reproduce artifacts or pursue infeasible targets. Inspired by how professional animators repair motion by preserving the intended action while correcting local artifacts, our key insight is that an imperfect reference can still convey reliable motion intent even when its exact joint trajectories are unsuitable for execution. We present **SetMimic**, a policy-in-the-loop interpretation framework that learns a humanoid policy for each imperfect retargeted motion by preserving reliable intent and resolving uncertain details through robot execution feedback. SetMimic estimates uncertainty to build a structured Motion Intent Representation, then samples and refines intent-consistent trajectories into a reference tube. Short policy rollouts and a learned feasibility critic guide refinement toward realizations that the robot can execute. An uncertainty-conditioned set-tracking objective tracks reliable components precisely while allowing uncertain components to vary within admissible regions, where physical objectives guide execution without a set-tracking penalty. This formulation couples motion interpretation with policy learning to preserve the intended action while accommodating imperfect video-derived references.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.