acceptodds
Under review as a conference paper at ICLR 2027

Kimodo-Policy: From Text-to-Motion Generators to Humanoid Vision-Language-Action Policies

Abstract

Humanoid Vision-Language-Action (VLA) models translate multimodal task understanding into coordinated whole body behavior. Existing approaches equip pretrained vision language or video generation backbones with action experts, but their pretraining does not directly establish whole body motion priors. Unified or upper-lower-body action designs also lack dedicated root, body, and hand modeling despite their distinct requirements. These limitations increase the burden of learning precise, coordinated actions from limited task-specific humanoid demonstrations. We introduce Kimodo-Policy, which initializes visually grounded control with pretrained text-to-motion (T2M) priors over inter-joint coordination and temporal structure. Its three stage root-body-hand architecture combines dedicated generators with explicit dependencies: root trajectories condition body generation, while future body features and visual observations guide hand control. We pretrain the policy on diverse humanoid demonstrations, conditioning actions on visual observations, language instructions, and motion history. With 678M total and only 92M trainable parameters, Kimodo-Policy achieves strong performance on HumanoidArena, SIMPLE, and real-world loco-manipulation tasks. Further experiments show performance gains from larger VLA pretraining datasets, greater model capacity, and stronger pretrained T2M priors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.