Dynamic Humanoid Fall Recovery: A Dataset and Robust Recovery Policy
Abstract
Humanoid fall recovery remains challenging due to the highly diverse and out-of-distribution recovery states from real-world falls, and the conflict between rapid physical recovery and stable, natural motion. Existing approaches mainly address this challenge through multi-stage curriculum learning. However, such pipelines require complex multi-stage configurations and suffer from policy degradation due to distribution shifts induced by rigid stage transitions. To address these limitations, we propose KiPAR, a recovery-oriented reinforcement learning framework that combines dynamic state initialization, history-consistent motion priors, and phase-aware reward scheduling. KiPAR bypasses canonical static initializations by using a comprehensive dataset of diverse, dynamic dangerous states as initialization source. To ensure coherence after resetting to these highly unstable states, we introduce kinematic prior history injection to reconstruct physically consistent observation histories. Furthermore, we introduce a phase-aware reward scheme that dynamically balances heuristic incentives and stabilization regularizers. Extensive experiments demonstrate that KiPAR achieves a 98.34% recovery success rate, a 2.77s average time-to-stand, and a low 3.12% recollapse rate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.