acceptodds
Under review as a conference paper at ICLR 2027

NaviVLA: Socially-aware Robot Navigation Vision-Language Action Model

Abstract

Socially-aware robot navigation remains meaningful yet challenging, as robots must jointly optimize goal reaching, collision avoidance, human comfort, and implicit social norms under uncertain environmental dynamics. Existing methods often depend on simplified 2D geometric states, while recent many VLA models are not directly designed for socially compliant navigation. We propose NaviVLA, an RL fine-tuned social navigation VLA planner equipped with a flow matching action expert. NaviVLA integrates BEV maps, multi-sensor input, and world model powered language context to perform social navigation reasoning. Moreover, to support scalable and realistic training, we further construct an IsaacSim-based social navigation benchmark with diverse pedestrian dynamics simulators, standardized baselines, and multiple metrics. Its crowd module integrates both rule-based models and diffusion-powered simulators trained on real-world crowd trajectory datasets. Additionally, we introduce a RLHF-based social norm encoder to align the policy with respect to human preferences. Experiments in simulation and a physical robot show that NaviVLA improves navigation success, safety, and social compliance, establishing a new VLA-based framework for socially-aware robot navigation. The project website is accessible at: https://sites.google.com/view/NaviVLA

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.