CoLM: Controllable Text-Driven Humanoid Loco-Manipulation Generation
Abstract
Humanoid loco-manipulation requires coordinated whole-body motion and accurate object interaction, yet collecting such demonstrations through teleoperation or motion capture is difficult to scale. Generative models provide a promising alternative, but language alone specifies only coarse task intent and offers limited control over object-level spatial goals. We present CoLM, a framework for generating and executing humanoid loco-manipulation from language and explicit object-level spatial constraints. Given a text description, object geometry, and the initial humanoid and object states, our model jointly generates humanoid and object motion while supporting spatial control ranging from a target object pose to a full trajectory. The generated motion is then executed by a motion tracking policy. We construct a humanoid loco-manipulation dataset by retargeting existing human-object interaction data to the Unitree G1 with contact-guided refinement. We further introduce the interaction field, a dense representation of humanoid-object spatial relations that serves as a shared interaction objective across generation and execution, being jointly predicted with motion during generation and used as a dense reward for tracking. Experiments show that our generator produces higher-quality motions than existing baselines and accurately satisfies object-level spatial constraints while preserving humanoid-object interaction, and our tracking policy achieves higher execution success. We further demonstrate the real-world applicability of CoLM through successful execution of diverse loco-manipulation tasks on the Unitree G1.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.