Semantics Within, Relations Between: Structuring Vision-Language Representations for Person Re-Identification
Abstract
Recent CLIP-based person re-identification (ReID) methods increasingly exploit language-aligned semantics through textual prompts, descriptions, or image-derived pseudo-words. However, image-derived semantics are typically constructed only after visual encoding, limiting their interaction with visual representation formation, while conventional sample-level objectives do not explicitly organize the relative ordering of hard competing identities. These limitations can leave fine-grained visual evidence underexploited and ambiguous identity neighborhoods insufficiently structured for retrieval. We present SWR, a framework that addresses these two limitations through Semantics Within and Relations Between. Its In-Encoder Visual Semantic Tokenization (IVST) introduces learnable semantic tokens into the visual Transformer, allowing them to evolve jointly with image patches, and collects their intermediate states across visual layers to construct image-conditioned pseudo-word representations. Its Identity-Relational Ranking (IRR) aggregates observations into identity prototypes and applies listwise supervision over hard identity neighborhoods in the fused pre-BN representation. On Market-1501 and MSMT17, SWR improves the PromptSG baseline by 2.0 and 1.7 mAP points, reaching 96.6% and 88.9% mAP, respectively. It further achieves 72.7% mAP on Occluded-Duke and 86.0% on CUHK03-D. Compared with IADT under matched training settings, SWR improves mAP by 0.4 and 0.5 points on Market-1501 and MSMT17 while using 29.1% fewer parameters and reducing training time by 29.6–33.6%. Controlled analyzes further validate multi-layer state selection, relational supervision location, hard-identity mining, and retrieval-boundary behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.