acceptodds
Under review as a conference paper at ICLR 2027

Beyond Global Conflict Minimization: Execution-Aware Skill Blending for Humanoid Control

Abstract

Composing pretrained skills offers a practical route to humanoid whole-body control, but disagreement between their action contributions is not itself a measure of execution quality. We study this distinction in a SkillBlender-based controller using mask-weighted action cancellation as a global conflict diagnostic. We introduce a jointwise execution-aware cost that couples bounded cancellation with excess joint tracking error and actuator saturation, and optimize it through PPO with a learned Lagrange multiplier. Across five paired Button training seeds, this formulation reduces mean tracking error, torque, clipping, and saturation relative to unconstrained blending, while increasing mean global conflict. Cost ablations show that minimizing cancellation alone does not reproduce these execution improvements. Re-evaluation of three Transfer training seeds at 80,000 iterations yields lower mean execution errors, but a controlled comparison of two fixed seed-33 policies over 20 matched evaluation episodes reveals substantially worse object placement and reversals in several execution metrics. These findings support separating conflict, physical execution, and task progress in the evaluation of skill composition. They do not establish universal superiority, constraint satisfaction, or real-robot safety. We provide an implementation audit and a reset-aware evaluation protocol to make these distinctions explicit.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.