acceptodds
Under review as a conference paper at ICLR 2027

DisCo: Disambiguating Multi-modal Floorplan Localization via SE(2)-Aware Contrastive Learning

Abstract

Visual Floorplan Localization (FLoc) aims to estimate camera poses by aligning visual observations with floorplans, yet it remains challenging due to structural aliasing in repetitive indoor environments. Such ambiguity causes distinct poses to exhibit similar visual-floorplan compatibility, resulting in multi-modal localization distributions with limited spatial separability and angular discriminability. Existing ray matching-based methods can generate geometry-consistent pose candidates, but lack effective mechanisms to disambiguate them. To address this issue, we propose DisCo, a visual-floorplan Contrastive Disambiguation framework for robust visual FLoc. First, we employ a depth-aware ray regression predictor to project monocular RGB observations into 2D ray primitives and generate multi-modal pose candidates through ray-floorplan matching. We further introduce a clustering-based hypotheses preparation strategy to extract representative FLoc hypotheses by exploiting the distribution of candidate poses. Second, we propose an -aware contrastive disambiguation strategy to learn fine-grained visual-floorplan compatibility beyond geometric matching. Through structured pose perturbations, DisCo improves local pose smoothness, spatial separability, and angular discriminability by capturing implicit structural semantics in local floorplan layouts without relaying on explicit semantic annotations. Extensive experiments on three challenging visual FLoc benchmarks demonstrate that DisCo consistently outperforms state-of-the-art methods and significantly improves the robustness and accuracy of visual FLoc.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.