Contrastive Reinforcement Learning over Adversarial Skill Priors for Long-Horizon Humanoid Navigation
Abstract
Long-horizon humanoid navigation couples dynamic balance with goal-directed exploration across large environments. Contrastive reinforcement learning (CRL) enables goal-conditioned exploration without reward engineering, but end-to-end training on a 28-DoF humanoid forces joint-space exploration, spending most of the sample budget discovering basic locomotion rather than navigating. Pretrained motor priors provide stable gaits from decoded latents, yet they are typically driven by policy-gradient methods that struggle with long horizons. We propose ASE-CRL, which pairs a frozen Adversarial Skill Embedding (ASE) prior learned from motion capture with a contrastive goal-conditioned learner. Neither tier requires reward shaping: ASE trains via adversarial imitation, and CRL learns purely through goal identification. The advantage of the contrastive objective over reward-driven learning grows with maze scale once exploration operates over skills rather than joint targets. ASE-CRL solves every maze layout in our suite, exceeding 97% success on the largest, whereas reward-driven and end-to-end baselines collapse as navigation horizons grow. On whole-body pushing, ASE-CRL is the only goal-conditioned method that transfers from character to object goals, where hindsight-relabeled SAC fails. Finally, we release HGC-NavBench, an Isaac Lab benchmark suite for online humanoid goal reaching.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.