STaG: Human-Annotation-Free Unsupervised Domain Adaptation for Text-Based Person Retrieval
Abstract
Real-world deployment of Text-Based Person Retrieval (TBPR) is bottlenecked by the absence of textual annotations in new domains. Existing Unsupervised Domain Adaptation (UDA-TBPR) methods still assume access to target-domain human-written annotations (albeit unpaired with images), limiting their practicality. We consider a more practical setting: human-annotation-free UDA-TBPR, with labeled image–text data in the source domain and annotation-free, unlabeled images in the target domain. This setting introduces two core challenges: the absence of cross-modal alignment in the target domain, and visual domain shifts from source to target. To tackle these, we propose STaG, a dual-path framework comprising Style-Agnostic Text Alignment (SATA) and Heterogeneous Graph Memory Alignment (HGMA). Specifically, SATA leverages source-domain data to decouple identity-discriminative semantics from machine-generated stylistic artifacts, establishing robust cross-modal alignment for the target domain. Meanwhile, HGMA exploits cross-domain semantic neighborhoods via a memory-based graph to learn domain-invariant visual features. Ultimately, STaG achieves state-of-the-art performance across three benchmarks under this realistic setting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.