Beyond Verbalized Feedback: Mining Intent-Aligned Visual Clues for Interactive Text-to-Image Person Re-Identification
Abstract
Text-to-image person re-identification (TIReID) aims to retrieve a target pedestrian from a large-scale image gallery using a natural-language description. Since a single description rarely captures the fine-grained details that distinguish visually similar pedestrians, interactive TIReID progressively completes the initial description through multi-round user feedback. However, existing methods rely on user feedback, and retrieved candidates are rarely used as a direct source of visual clues. As a result, subtle variations in color tone, texture, and clothing style, which are difficult to convey through textual feedback, can be easily overlooked. To address this issue, we propose Visual Clue-guided Interactive Retrieval (VCIR), a plug-and-play framework that mines intent-aligned visual clues from retrieved candidates. Specifically, we first design a Visual Clue Extraction (VCE) module that guides a Multi-modal Large Language Model (MLLM) to contrast these candidates against the user intent and extract discriminative clues. Then, we introduce a learnable retrieval token that distills these clues into a continuous representation, bypassing the information loss caused by verbalizing them into discrete text. We further embed VCE into a multi-round interactive retrieval framework, where the extracted clues directly serve as the retrieval query at each round and progressively steer retrieval toward the target pedestrian. VCIR introduces only 32.9M trainable parameters and easily integrates with existing interaction paradigms. Extensive experiments on Interactive-PEDES and MInterPEDES demonstrate that VCIR achieves state-of-the-art retrieval performance. The code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.