acceptodds
Under review as a conference paper at ICLR 2027

GUIDE: Geometry-Grounded Identity Structure Learning for Multi-Modal Object Re-Identification

Abstract

Multi-modal object re-identification (ReID) aims to retrieve the same target across cameras by exploiting complementary cues from heterogeneous modalities. Recent graph-based approaches have further advanced multi-modal ReID by explicitly modeling relational dependencies among local regions and semantic representations. However, graph-based methods typically derive relational topology from feature or semantic similarity, causing corresponding regions across modalities to aggregate different contextual information during relational propagation, thereby weakening the consistency of identity-related structural cues. To address this issue, we propose GUIDE, a Geometry-Grounded Identity Structure Learning framework for multi-modal object ReID. GUIDE decouples relational topology from modality-dependent appearance by grounding connectivity in patch geometry while preserving modality-specific representations. Specifically, the Grid Graph Encoder (GGE) restores patch tokens to their native two-dimensional layouts and propagates local identity cues over geometry-defined neighborhoods using a graph encoder shared across modalities. The Axial Strip Readout (ASR) preserves spatial composition by aggregating graph-refined tokens into ordered strips with learnable queries. Together, GGE and ASR produce structured identity representations that complement global semantic features for retrieval. Extensive experiments on RGBNT201, RGBNT100, and MSVR310 demonstrate the effectiveness of GUIDE across person and vehicle ReID benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.