acceptodds
Under review as a conference paper at ICLR 2027

HoRD: Robust Humanoid Control via History-Conditioned Reinforcement Learning and Online Distillation

Abstract

Humanoid robots hold great promise for general-purpose embodied intelligence, but their control policies often suffer significant performance drops under small changes in robot and environmental dynamics. Dense full-body references simplify tracking but impose a restrictive deployment command, whereas sparse targets leave uncommanded joints underdetermined. We present HoRD, a two-stage learning framework for whole-body humanoid control under partial observability and varying dynamics. A teacher first learns dense-reference tracking under episode-level domain randomization. Its control behavior is then transferred through online distillation to a student that receives partial observations and Standardized Sparse-Joint Representation (SSJR) commands over a fixed set of future key-joint targets. Both policies use a History-Conditioned Dynamics Representation (HCDR), a history encoder that summarizes recent state–action interactions into a control-oriented representation. HoRD achieves State-Of-The-Art (SOTA) performance across four conditions spanning Isaac Lab and Genesis, including zero-shot cross-simulator transfer under fixed and randomized dynamics. Ablations show that episode-level domain randomization and HCDR play complementary roles under dynamics shift. Qualitative Unitree G1 rollouts further illustrate real-world deployment without fine-tuning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.