Humanoid Motion Foundation Model for Multimodal Streaming Closed-loop Interaction
Abstract
Recent advances in vision-language-action (VLA) systems have enabled increasingly general-purpose embodied behaviors. However, existing approaches still formulate humanoid interaction as a static conditional generation problem, limiting their ability to support temporally evolving multimodal interaction and execution-aware behavioral adaptation in open-world environments. In this paper, we propose MSC, a multimodal streaming behavioral framework for closed-loop humanoid interaction. Instead of conditioning humanoid behaviors on fixed multimodal conditions, MSC reformulates humanoid control as a streaming conditional flow process under continuously evolving multimodal and execution-aware states. Built upon a structured conditioning representation and a streaming flow matching mechanism, MSC enables continuous interaction and closed-loop behavioral refinement during execution. Experiments show that MSC consistently improves motion quality, multimodal alignment, and robustness across text-, vision-, and audio-conditioned humanoid control tasks. We further demonstrate real-world deployment on Unitree G1 humanoids under interleaved multimodal interactions, validating the practicality of streaming closed-loop humanoid interaction in real-world environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.