Dynamic Audio-Visual Navigation in Continuous Environments with Moving Targets
Abstract
In real-world scenarios, sound-emitting entities such as people, pets, and mobile robots may move continuously or emit sound intermittently, and their sounds may be mixed with those from irrelevant sound sources. While recent studies have extended audio-visual navigation from discrete, non-semantic settings to continuous, semantic settings, these advances assume stationary targets. We therefore introduce dynamic audio-visual navigation in continuous environments (DANCE), a new task in which an embodied agent uses audio-visual observations to navigate toward a single sound-emitting target that may be moving or stationary. Navigating toward a moving target is challenging because previously observed spatial cues can quickly become outdated as the target moves. To address this challenge, we propose DyGMap, which couples audio-visual goal estimation with dynamic goal mapping. The former integrates complementary acoustic and visual evidence to estimate the target's semantic category and location, while the latter projects these estimates onto category-specific goal maps and decays outdated or spurious target evidence. Extensive experiments in clean and noisy settings demonstrate that DyGMap improves the success rate by approximately 11.6 percentage points over the strongest evaluated baseline, demonstrating the importance of robust goal reasoning and dynamic goal mapping for efficient navigation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.