Interactive Semantic Verification for Unsupervised Domain Adaptation in Text-to-Image Person Re-Identification
Abstract
Unsupervised domain adaptation (UDA) has received increasing attention in text-to-image person re-identification (TIReID), as it enables models to adapt to an unlabeled target domain without costly annotations. However, under domain shift, existing feature-based adaptation methods face reliability challenges in both target domain pseudo correspondence construction and cross-domain representation learning. To this end, we propose an Interactive Semantic Verification framework (ISV) that leverages a multimodal large language model (MLLM). Specifically, to alleviate noisy pseudo image-text correspondences in the target domain caused by unreliable feature similarity, we develop Interactive Multi-evidence Learning (IML). It first performs balanced reciprocal matching to identify candidate correspondences and then employs an MLLM for semantic verification, leveraging external semantic knowledge to compensate for unreliable feature evidence, thereby removing mismatched pairs and recovering missed positive matches. Moreover, to improve cross-domain representation learning, we introduce Density-filtered Intermediate-domain Learning (DIL), which identifies and filters extreme feature outliers in low-density regions and explicitly models the difference between source and target features to construct more reliable intermediate representations. Extensive experiments on three TIReID benchmarks demonstrate that ISV consistently improves cross-domain retrieval performance and generalizes well across different source models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.