IndustryNav: Evaluating Spatial Reasoning and Safety of Embodied Agents in Dynamic Industrial Navigation
Abstract
Visual large language models (VLLMs) are increasingly used as embodied agents, yet industrial warehouses demand more than reaching a goal: agents must navigate safely around moving workers and vehicles. To evaluate embodied agents' goal reaching and safety, we introduce IndustryNav, a simulation environment with 24 warehouse scenes for closed-loop navigation to specified destinations. Alongside task success and efficiency, we introduce safety-aware metrics that capture collisions and near-collision warnings along each trajectory. Experiments with eleven leading open- and closed-source VLLMs and three navigation baselines show that agents can reach their goals yet collide with moving obstacles, exposing a gap between success and safe navigation. This gap matters in safety-critical settings, where a single collision can end the task. IndustryNav also supports policy learning through automated oracle trajectories and simulator interaction, with held-out scenes for generalization tests. These experiments show that policies can succeed on new routes in familiar warehouses yet fail to transfer to held-out layouts, making cross-scene generalization an open challenge.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.