acceptodds
Under review as a conference paper at ICLR 2027

EgoExo-VLN: Egocentric-Exocentric Vision-and-Language Navigation in Dynamic Environments

Abstract

Existing Vision-Language Navigation (VLN) methods primarily rely on egocentric observations, limiting their ability to perceive important environmental changes beyond the robot's current view. To address this limitation, we propose a novel VLN task, named Egocentric-Exocentric Vision-Language Navigation (EgoExo-VLN), which enables agents to follow language instructions in dynamic indoor environments by combining egocentric observations with information from fixed exocentric cameras. EgoExo-VLN goes beyond simply providing additional views by requiring agents to identify task-relevant information across viewpoints, determine whether observed evidence remains valid over time, and decide whether it should influence the current VLN decision. To systematically evaluate these capabilities in the EgoExo-VLN setting, we construct the first VLN benchmark that explicitly incorporates exocentric cameras, consisting of 8,096 episodes across 90 indoor scenes with scene-disjoint training, validation, and test splits. The benchmark focuses on dynamic target relocation and introduces controlled absent, conflict, and non-conflict pedestrian conditions, while also including static-scene tasks to cover common VLN scenarios. Building on this benchmark, we propose a VLN framework that selects task-relevant camera information, updates target representations over time to revise VLN goals, and reasons about pedestrian conflicts to determine when the agent should wait or continue moving. EgoExo-VLN establishes a new benchmark for studying VLN with egocentric observations and exocentric camera information in dynamic environments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.