StaCo: Streaming Text-Driven Humanoid Control with Shared State–Command Latents
Abstract
Existing latent-driven methods for text-driven humanoid control streamline the generation-to-control pipeline, but typically learn motion representations independently of control policies and rely on command-centric latents. This separation yields representations that are weakly aligned with robot execution and prevents the generator from incorporating the robot's actual execution state into future predictions. We propose StaCo, a shared state–command latent framework that couples motion representation learning with humanoid control and closes the feedback loop between streaming generation and physical execution. StaCo trains a Control-Native Motion VAE through teacher-student policy distillation, yielding a robot-native latent space that serves both as an execution-state representation and as an executable control target. Built on this shared latent space, a streaming DiT predicts future command latents from language and historical execution latents, while proprioceptive feedback is re-encoded after execution to inform subsequent generation. Experiments demonstrate that the learned latent space supports robust humanoid control while retaining high-fidelity motion structure, and that the streaming generator produces high-quality motions consistent with language instructions. End-to-end evaluation achieves a 97.35% strict control success rate in MuJoCo, while text-conditioned generation attains an FID of 19.14 on HumanML3D-derived robot motions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.