acceptodds
Under review as a conference paper at ICLR 2027

Knowing Where to Go: Generalizing Mobile Manipulation to Unseen Environments from a Single Video

Abstract

Vision-language-action policies have achieved impressive mobile manipulation capabilities, but remain brittle in unfamiliar multi-room environments. We find that much of this brittleness stems from failures of spatial generalization: strong policies often resort to reactive exploration when task-relevant locations move outside familiar layouts. Oracle spatial guidance largely prevents this degradation. We study how spatial context from a prior environment walkthrough video should be represented to guide mobile manipulation policies under environment shift. We find that strong in-distribution spatial prediction can be a poor predictor of out-of-distribution transfer, while models that explicitly encode 3D structure generalize more robustly to new robot environments. Building on this, we develop a hierarchical policy that uses a single environment video to provide spatially grounded guidance to a VLA. We find that navigation performance of mobile manipulation policies is heavily bottlenecked by spatial understanding, and that our strongest policy more than doubles the placement rate of its flat counterpart ( vs. ) in our hardest setup. For more information, see https://know-where-to-go.github.io/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.