acceptodds
Under review as a conference paper at ICLR 2027

Aligning without Homogenizing: Jointly Explainable Common-Specific Decomposition for Multi-Modal Object Re-Identification

Abstract

Multi-modal object Re-Identification aims to reduce cross-modal discrepancies while preserving identity-relevant modality-specific cues. Previous methods often align representations from heterogeneous modalities uniformly; however, treating common and modality-specific information alike may over-smooth useful modality-specific cues and cause information loss. To address this issue, we propose JESD-ReID, which first distinguishes these two types of evidence using a joint-explainability criterion and then coordinates each according to its role. Jointly Explainable Semantic Decomposition (JESD) uses shared semantic queries to form compact observations with shared query-slot indexing and reconstructs them jointly from a single latent code through constrained modality-conditioned dictionaries. The reconstructed evidence forms the common components, while the remaining token-level evidence is retained as modality-specific residuals. Low-volume regularization organizes the common span, whereas relational optimal transport coordinates local relational structure in the residuals without forcing feature values to coincide. Six lightweight modality–role experts use sample-dependent gates to control how strongly each branch contributes, while a jointly predicted shared grid controls where complementary evidence is sampled for aggregation. JESD-ReID achieves 82.4%, 88.0%, and 66.6% mAP on RGBNT201, RGBNT100, and MSVR310, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.