acceptodds
Under review as a conference paper at ICLR 2027

RePO: Reference-Relative Policy Optimization for High-Fidelity Humanoid Motion Imitation

Abstract

Physics-based humanoid controllers trained with Proximal Policy Optimization (PPO) have achieved high success rates in tracking diverse reference motions. However, successful execution does not necessarily imply faithful imitation: policies can increase cumulative rewards by surviving longer while adopting safer behaviors that deviate from the reference. To address this mismatch, we propose Reference-Relative Policy Optimization (RePO), which combines a critic-based viability advantage to support robust, long-horizon physical execution with a group-relative fidelity advantage to encourage faithful reference tracking. Instead of assigning a coarse trajectory-level score to each rollout, RePO normalizes returns-to-go across executions of the same motion at each aligned reference timestep, yielding dense, step-wise group-relative advantages that are blended with the viability advantage for policy updates. To reduce interference from survival differences, we introduce a survival gate that activates the fidelity advantage only when group-level execution is sufficiently reliable. We further adopt a kinematic cold start to initialize the policy from motion data, providing a reference-following prior for early group comparisons. Experiments on AMASS demonstrate improved motion fidelity and execution success without PHC's progressive training strategy. Furthermore, we demonstrate that our approach can facilitate the learning of complex motions. We validate this capability through challenging human–object interaction experiments on OMOMO, where RePO enables direct full-dataset training while the InterMimic fails to learn an effective policy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.