SocioWM: Social-Conditioned World Model for Social Context Vision-and-Language Navigation
Abstract
Conventional vision-and-language navigation (VLN) enables embodied agents to navigate static scenes by following natural-language instructions. Existing Social-VLN works further introduce moving pedestrians into navigation scenarios, but primarily model them as dynamic obstacles or generic visual entities, overlooking the social context that influences pedestrian behavior, including personal attributes and human–robot interactions. Consequently, existing methods may fail to account for attribute-dependent differences among pedestrians and their interaction responses to robot actions, leading to navigation behavior that prioritizes goal progress while neglecting social norms. To address this limitation, we formulate Social Context Vision-and-Language Navigation (Social Context-VLN), a new VLN setting that incorporates pedestrian motion and social context into navigation decision-making. To support this setting, we construct , the first benchmark for Social Context-VLN. It explicitly annotates Pedestrian Attribute, including age and gender, and Human–Robot Interaction as social context, and contains more than 35K language-guided navigation episodes involving 1.4K pedestrians. To validate the effectiveness of these structured annotations, we further introduce , a socially conditioned world model that predicts the social consequences of candidate routes and selectively modulates candidate world states for action selection. Experiments on show that improves the social success rate (SR) from 18.27% to 27.07%, while reducing collisions and improving personal-space compliance. Evaluation of existing VLN methods further shows that strong conventional navigation performance does not necessarily imply socially compliant behavior. Real‑world experiments verify our method transfers to unseen indoor environments and real‑world pedestrian interactions.Ours project page:https://anonymous.4open.science/r/iclr2027anony
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.