BGDL: Learning Whole-Body Humanoid Control from Demonstrations with Physical Action Geometry
Abstract
Demonstrations are useful for whole-body humanoid control, but their value is not uniform along a trajectory. In permissive states, many actions can still complete the task, whereas near foothold limits, balance boundaries, or required contacts, small deviations can change the outcome. We introduce **Boundary Guided Demonstration Learning (BGDL)**, which uses this variation in local action freedom to determine where demonstration guidance should be strongest. BGDL derives state-dependent guidance from physical margins, proximity to required interactions, and near-term failure risk, strengthening imitation where action choices are more restricted while leaving greater freedom for reinforcement learning elsewhere. We further characterize how allowable deviation from a demonstrated action depends on margin reserve, local action sensitivity, and margin-estimation error, giving a local sufficient condition for preserving a one-step physical-margin condition. We evaluate BGDL on five whole-body humanoid tasks: stair climbing, curb traversal, ground object pick-up, chair sitting, and object carrying, and transfer the learned policies to a Unitree G1 without policy fine-tuning. BGDL improves average success from 86.0% to 91.2% over reference tracking while reducing falls from 12.8% to 8.0%. With matched demonstration data, total imitation weight, and environment interactions, BGDL improves success by 3.4 percentage points over random assignment of the same guidance weights. This suggests that the placement of demonstration supervision matters beyond its total amount.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.