acceptodds
Under review as a conference paper at ICLR 2027

DreamPolicy: Scaling Humanoid Locomotion through Terrain-Aware Generative Planning

Abstract

Achieving versatile humanoid locomotion with a single policy remains a key scalability challenge. Existing approaches often distill multiple terrain-specific policies into a unified policy, but such distillation mainly transfers individual locomotion primitives and provides limited ability to compose them when encountering novel and structurally different terrains. We present DreamPolicy, a scalable framework for humanoid locomotion through terrain-aware generative planning. We first construct a modular multi-terrain benchmark with over a dozen primitive terrains and their combinations, and collect 75M transitions of expert rollouts from specialized policies, providing diverse locomotion primitives and terrain-conditioned behaviors for learning. On this dataset, we train a terrain-aware autoregressive diffusion planner to generate a humanoid motion imagination (HMI), a physically plausible future state trajectory conditioned on state history and terrain information. A unified HMI-conditioned policy then tracks these generated future states, allowing locomotion behaviors to be dynamically composed according to the encountered terrain without terrain-specific reward design. DreamPolicy scales consistently with offline data and generalizes beyond the terrains covered by the source policies. Experiments show substantial improvements over strong baselines on unseen and composite terrains, together with zero-shot transfer to a Unitree G1. These results demonstrate that terrain-aware generative planning provides a scalable approach to humanoid locomotion beyond the "one task, one policy" regime.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.