Scaling Object Navigation With Cross-Domain Reinforcement Learning
Abstract
Object-goal navigation (ObjectNav) requires an agent to locate a specified object in an unfamiliar environment. Policies trained on limited 3D scenes generalize poorly to unseen layouts, yet collecting and annotating large-scale 3D training data remains expensive. To address this limitation, we propose a cross-domain RL framework that uses 2D semantic-map environments as an additional source of interactive training experience for scaling ObjectNav. First, TensorEnv turns static semantic maps into interactive 2D environments. Second, ParHab accelerates Habitat sampling and map construction. Third, MapAct takes partial semantic maps as input and predicts navigation actions. Fourth, NavPRL bridges the domain gap by transferring knowledge from a teacher model trained in TensorEnv to a student model trained in ParHab. Experiments show that incorporating 2D semantic maps improves navigation success and that increasing the number of training floorplans yields further gains at fixed policy capacity and per-stage budgets. The 12.08M-parameter MapAct supports local deployment on a 4 GB NVIDIA Jetson Nano and achieves 68.77%/37.60% SR/SPL on HM3D and 44.19%/19.57% on MP3D, competitive with methods based on large language and vision-language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.