Beyond Ranked Gallery: Conformal Identity Set for Text-Based Person Retrieval
Abstract
Text-based person retrieval enables users to search for individuals in an image gallery using natural-language descriptions. Existing methods primarily focus on cross-modal matching and ranking accuracy, whereas practical retrieval also requires deciding how many candidate identities to return. A fixed cutoff may exclude the target for ambiguous descriptions or return unnecessary candidates for distinctive ones. Conformal prediction offers a principled framework for constructing such sets by calibrating candidate-level scores. A natural approach is to derive these scores from softmax-normalized retrieval similarities. However, this normalization introduces a temperature dependence that can alter set sizes and coverage evenness without changing the retrieval ranking. To understand this dependence, we analyze the low-temperature behavior of representative softmax-based conformal scores. We prove that, after order-preserving transformations, these scores converge to a common similarity-margin score as temperature approaches zero. Motivated by this shared limit, we introduce a temperature-free method for constructing Conformal Identity Sets through direct margin calibration. Specifically, our method retains identities within a common calibrated similarity gap of the top-ranked identity, allowing set sizes to adapt to each query's similarity profile. The resulting sets provide finite-sample marginal coverage at a user-specified level under exchangeability, without softmax, temperature tuning, or retraining. Experiments with five frozen retrievers across three benchmarks demonstrate empirical coverage close to the target and a favorable trade-off between set size and coverage evenness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.