acceptodds
Under review as a conference paper at ICLR 2027

HUMANO: A Large-Scale Benchmark for Generalist Humanoid Loco-Manipulation

Abstract

Recent advances in vision-language-action models and whole-body control have rapidly expanded the capabilities of humanoid loco-manipulation, enabling generalist policies to perform increasingly diverse tasks with a single model. However, existing benchmarks often focus on individual tasks or small task suites, making it difficult to systematically evaluate multi-task capabilities and generalization. To fill this gap, we present HUMANO, a large-scale benchmark for generalist humanoid loco-manipulation. HUMANO contains 100 tasks organized into eight complementary suites, spanning 90 interactive scenes and over 2K object assets, with 5.5K in-distribution whole-body demonstrations. The benchmark covers fundamental skills, goal-directed loco-manipulation, generalization across spatial configurations, objects, robot configurations, and goals, as well as long-horizon skill composition and state-dependent reasoning. HUMANO further provides seven controlled OOD settings covering visual, linguistic, and robot-initialization shifts, together with 27.5K OOD demonstrations for future training and adaptation studies. All policies operate through a shared SONIC-based motion-latent interface, enabling standardized comparison under a common whole-body execution framework. Experiments with representative imitation and vision-language-action policies demonstrate the emerging potential of generalist models for multi-task humanoid loco-manipulation. Although multi-task learning remains more challenging than single-task training, performance continues to improve with increasing demonstration scale. Preliminary sim-to-real results further show that policies trained on HUMANO simulation data can transfer executable whole-body behaviors to a physical humanoid robot.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.