acceptodds
Under review as a conference paper at ICLR 2027

MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining

Abstract

Mobile manipulation lets a robot control the extent of its own workspace, but it couples two challenges that fixed-base policies can largely avoid: perception must remain spatially grounded under continuous ego-motion, and heterogeneous arm and base actions must be coordinated. We present Goku, a vision-language-action model built around seeing, coordinating, and imagining arm-base collaboration. Goku conditions its action expert on sparse multilevel VLM features, supervises the shared perceptual representation with a training-only branch that predicts future features from a frozen geometric teacher, and generates manipulation and body actions with MM-APT, a two-stream transformer that couples the streams through masked near-far joint attention and directly predicts clean action chunks. A toy experiment motivates clean-action prediction for coupled action streams. In the full policy, replacing clean-action prediction with velocity prediction reduces success on RoboCasa365 composite-seen tasks from 32.8% to 29.2%. We pretrain Goku on more than 5,000 hours of robot data spanning over 400K episodes, 12 datasets, and 17 embodiments, including MM-30, a new multi-view dual-arm mobile manipulation dataset. Goku achieves 44.71% success on EBench, 61.2% on RoboCasa365, the highest mean success across three ManiSkill-HAB suites, 99.1% on LIBERO, 82.8% on LIBERO-Plus without perturbation training, and 83% on five long-horizon real-world tasks, 12 percentage points above pi_0.5.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.