Foot Height Symmetry and Torso Roll Reward Shaping for Robust Upright Standing Balance and Posture Stabilization in High-DoF Humanoids
Abstract
Training bipedal upright balance on MuJoCo Humanoid-v5 (348 observation dims, 17 actuated DoFs) with standard PPO is notoriously unstable. Value function loss diverges wildly, and the agent frequently cheats the objective by leaning onto one leg before toppling over in fewer than 100 steps. We resolve these bottlenecks with a physically grounded reward setup paired with stable training mechanics. Specifically, we introduce a bilateral foot-height symmetry penalty (Rfeet) alongside quaternion-derived torso roll leveling (Rroll) and target-height Gaussian shaping (Rheight centered at z = 1.35m). Running this on a 256 x 256 actor-critic architecture with VecNormalize and value-loss clipping (eps_vf = 0.2) effectively eliminates the degenerate single-leg bias—slashing it from 88.2% down to 1.4%. The policy bounds torso roll strictly within +/-1.2 deg and sustains upright balance beyond 850 steps, delivering a 12.5x survival gain over baseline PPO.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.