: A Steerable Humanoid VLA Model for Long-Horizon Loco-Manipulation
Abstract
Vision-language-action models predict action chunks efficiently but tend to blindly follow training trajectories when facing out-of-domain tasks. While world models demonstrate more robust generalization, they require longer inference times, which critically hinder humanoid loco-manipulation, where real-time execution is demanded. Combining the best of both, we present , a hybrid humanoid foundation model with unprecedented steerability that enables up to five minutes of autonomous execution on long-horizon, whole-body loco-manipulation tasks. Specifically, we first train a predictor: a time-conditioned image world model that predicts a future goal image to steer task execution, conditioned on a time horizon and a language instruction. We then re-purpose an existing VLA model, which predicts actions from solely the current visual observation, to additionally condition on a goal image. It acts as a planner in action space, outputting actions that steer the humanoid from the current observation toward the goal image. Finally, we implement a real-time execution pipeline that streams actions to a whole-body controller, achieving whole-body control capable of more diverse motions, such as kicking and bending, that are challenging for a decoupled controller. To obtain steerability, we draw on priors from the existing image generation model and further condition it on the time horizon and language instruction. In addition, we pretrain the VLA model with variable-length action chunk prediction, allowing it to develop a better understanding the dynamics between the current observation and goal image. Compared with baseline methods, , achieves better generalization, more reliable language instruction following, and higher success rates on long-horizon tasks. Through exhaustive experiments across seven real-world tasks, we demonstrate that our model achieves the highest overall performance, outperforming the strongest baseline by 25% in success rate in novel scenes with unseen objects and layouts, and by 43% in task progress on cross-task transitions, supporting whole-body loco-manipulation tasks lasting up to five minutes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.