CTRL: Consistency-Aware Text-Driven Representation Learning for Multimodal Vehicle Re-Identification
Abstract
Multimodal vehicle Re-Identification (ReID) aims to retrieve the same vehicle by utilizing complementary information across heterogeneous modalities. Recent studies further generate corresponding vehicle descriptions for images in each modality and leverage these textual semantics to guide visual feature learning and supply identity information. However, existing generation pipelines often emphasize descriptive completeness while overlooking the distinct requirements of text as visual guidance and as identity supplementation. Moreover, prior methods provide limited control over how textual semantics guide visual feature formation during encoding, and lack selective supplementation for each modality based on complementary identity cues retained by other modalities. To address these issues, we propose CTRL, a Consistency-Aware Text-Driven Representation Learning framework. Its text generation pipeline unifies attribute values and models modality-specific visibility to generate descriptions supported by the corresponding images and consistent in attribute values across modalities. Intra-modal Text-guided Asymmetric Propagation (ITAP) uses text to regulate information propagation during visual encoding, limits the information that identity-discriminative features aggregate from regions with weak identity information, and updates these regions in a controlled manner. Cross-modal Text-guided Semantic Completion (CTSC) supplements each modality with cues from other modality descriptions that are selected as receiving sufficient source-visual support while being insufficiently represented in the target text and visual features. Experiments on three multimodal vehicle ReID datasets show that CTRL significantly outperforms existing methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.