WorldPixelNav: Deployment-Oriented Pixel-Goal Navigation with Future-State-Aware Trajectory Generation and Hierarchical Correction
Abstract
Pixel-goal navigation bridges high-level reasoning and low-level control in vision-language navigation by selecting an image pixel as the next local target. Existing imitation-based planners, however, treat the pixel goal as a static target and do not model the spatial consequences of their motion, leading to trajectory drift and collision. We propose WorldPixelNav, a deployment-oriented pixel-goal navigation framework with future-state-aware trajectory generation and hierarchical correction. WorldPixelNav converts the pixel goal into a robot-centric planar coordinate, encodes depth observations, motion history, this coordinate, and a fixed instruction into shared conditional memory, predicts a future spatial latent from this memory, and decodes the latent into continuous waypoints. Training combines trajectory imitation, future-state prediction, and deployment-oriented supervision. During execution, safety verification triggers trajectory reconnection, in-place re-observation, or historical-viewpoint recovery without requesting a new high-level pixel goal. WorldPixelNav achieves 96.48% success rate and 0.07% collision-episode rate in unseen simulated scenes, and 83.3% and 6.7% in unseen real-world scenes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.