HearVGo: Benchmarking Sound-Conditioned Trajectory Adaptation For Urban Walking
Abstract
Environmental sound can reveal nearby activity and potential hazards along a route. Yet outdoor navigation largely centers on vision, while audio-visual navigation often treats sound as a cue to the destination. How sound should instead guide route adjustments toward an independent goal remains underexplored. We introduce HearVGo, a benchmark that tests this capability in omni models using brief multimodal observations. We augment CARLA with 1,288 environmental sound clips rendered as spatial audio and collect human walking demonstrations. The resulting 536 trajectory instances ask models to Adapt the path when sound warrants a change or Preserve the path otherwise. We evaluate 14 omni models under a common zero-shot protocol and complement the main comparison with controlled audio and vision ablations. These results show model-dependent gains from audio and few reversals of predicted avoidance direction when the left and right audio channels are exchanged. Some models retain substantial performance using only motion and the goal. Trajectory diagnostics show that goal progress can mask poorly formed waypoint sequences. Together, these results expose limits in the evaluated omni models' ability to determine when sound should alter a route and to express that decision as a coherent local trajectory. Code and data are publicly available: https://anonymous.4open.science/r/HearVGo.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.