AVLoc: Robust Audio-Visual Floorplan Localization under Sensor Noise and Environmental Change
Abstract
Floorplan localization estimates an agent's pose within a 2D floorplan from egocentric observations. Existing methods draw every prediction from a single sensor, a camera, whose limited field of view and vulnerability to dim corridors, motion blur, sensor noise, and occluded walls mean that one corrupted image corrupts every cue the model has. We present AVLoc, which pairs vision with binaural acoustic echoes. Echoes are governed by geometry rather than appearance, they cover the full , and are largely unaffected by the conditions that disable a camera. Complementarity alone does not give robustness, since a joint encoder can propagate corruption from one modality into the other. AVLoc therefore combines unimodal and multimodal experts, each predicting the surrounding layout with its uncertainty. The router weights the experts by these uncertainty estimates, so the model dynamically shifts its reliance toward whichever modality the current observation supports, and the fused uncertainty carries through to floorplan matching. % rather than being discarded at fusion. Under maximum visual corruption, AVLoc retains recall@1m against for the strongest vision-only baseline, and when vision and audio are degraded at once. Because echoes also constrain geometry beyond the field of view, AVLoc improves clean-condition accuracy as well, raising recall@1m from to on Gibson(f) and from to on HM3D. Code and benchmarks will be released upon publication.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.