GUIDE: Geometry-Grounded Identity Structure Learning for Multi-Modal Object Re-Identification
Abstract
Multi-modal object re-identification (ReID) aims to retrieve the same target across cameras by exploiting complementary cues from heterogeneous modalities. Recent graph-based approaches have further advanced multi-modal ReID by explicitly modeling relational dependencies among local regions and semantic representations. However, graph-based methods typically derive relational topology from feature or semantic similarity, causing corresponding regions across modalities to aggregate different contextual information during relational propagation, thereby weakening the consistency of identity-related structural cues. To address this issue, we propose GUIDE, a Geometry-Grounded Identity Structure Learning framework for multi-modal object ReID. GUIDE decouples relational topology from modality-dependent appearance by grounding connectivity in patch geometry while preserving modality-specific representations. Specifically, the Grid Graph Encoder (GGE) restores patch tokens to their native two-dimensional layouts and propagates local identity cues over geometry-defined neighborhoods using a graph encoder shared across modalities. The Axial Strip Readout (ASR) preserves spatial composition by aggregating graph-refined tokens into ordered strips with learnable queries. Together, GGE and ASR produce structured identity representations that complement global semantic features for retrieval. Extensive experiments on RGBNT201, RGBNT100, and MSVR310 demonstrate the effectiveness of GUIDE across person and vehicle ReID benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.