acceptodds
Under review as a conference paper at ICLR 2027

Language-Grounded Controller Latents for Humanoid Whole-Body Control

Abstract

Language-conditioned humanoid motion is commonly built as a cascade in which a text-to-motion model generates a kinematic trajectory that is subsequently tracked by a separate controller. We instead generate the latent commands of a frozen whole-body controller directly from text, removing runtime retargeting and motion tracking from the primary generation path. We also incorporate differentiable kinematics into diffusion sampling to impose hand and foot target constraints directly in the controller latent space, without retraining the generator or controller. Experiments demonstrate strong language alignment, high completion rates, and effective spatial-constraint preservation after physical execution. We further demonstrate diverse behaviors in simulation and on the real Unitree G1. Beyond generation and execution, we characterize the semantic and physical structure of the frozen controller latent representation and examine whether this structure is preserved in latents generated directly from language. Our analysis shows that the 64-dimensional latent representation encodes joint configuration, locomotion direction and intensity, and gait phase in a distributed manner, with much of this structure preserved after finite scalar quantization and also present in the generated latents. We further show that text-conditioned motion characteristics can be composed in latent space to generate motions combining multiple attributes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.