Where to Stand, What to Grasp: Occupancy-Aware Coupling of Mobility and Manipulation
Abstract
Mobile manipulation is challenging not because navigation and manipulation are individually difficult, but because they are tightly coupled: the robot’s base pose determines whether the arm can reach the target. Yet prevailing vision-language-action (VLA) models largely collapse this coupling into a monolithic action head, treating robot pose as merely another concatenated feature. We present M-VLA, which makes three modifications to the standard trunk- plus-flow-expert architecture. First, a Grounded Geometric Tokenizer (GGT) enables geometry-aware representation by converting the SLAM occupancy map and robot poses into metric-scale-preserving tokens within a shared physical coordinate frame, allowing map geometry and robot localization to interact through joint attention in the backbone. Second, Localization-as-Operator Fusion represents robot localization not as a conventional feature but as an SE(2) transformation operator that transports intent between the navigation (world) and manipulation (body) frames. It also incorporates an uncertainty-derived trust scalar distilled from the otherwise discarded SLAM covariance. Third, Dual Functional Experts replace the single action head with separate navigation and manipulation flow-matching experts coupled by a Coordination Bridge, enabling each expert to pretrain on abundant single-domain data while exchanging task intent at every layer. On a large- scale, long-horizon mobile-manipulation benchmark comprising tasks lasting from one to more than fourteen minutes and requiring cross-room navigation across multiple indoor scenes, \texM-VLA} achieves the strongest overall performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.