Bounded Contextual Refinement with Neighborhood Support for Text-Aerial Person Retrieval
Abstract
Text-aerial person retrieval aims to retrieve images of a target person from an aerial gallery given a natural language description. Variations in viewing angle, altitude, and occlusion can obscure the appearance cues described in the query, making visually similar individuals difficult to distinguish. Existing methods improve image-text alignment, but scoring each gallery image independently does not fully exploit context from other gallery images. In this paper, we propose Bounded Contextual Refinement (BCR), which improves aerial representations and retrieval rankings through bounded feature calibration and reciprocal gallery context. Bounded Feature Calibration (BFC) learns a norm-constrained residual transformation by contrasting the highest-scoring positive and negative gallery images for each text query. The constraint limits the magnitude of each feature correction while allowing the representation to adapt. Reciprocal Contextual Ranking (RCR) addresses two distinct retrieval decisions. It constructs candidate context vectors from reciprocal neighborhoods and learns bounded score corrections for candidate selection among the initial top five candidates. It then combines matching scores with neighborhood support to rank the remaining candidates within fixed intervals. To strengthen supervision for aerial appearance, Auxiliary Description Supervision (ADS) supplements training of the CLIP encoder with descriptions generated from aerial training images and reviewed against them by the same frozen vision-language model. Experiments on AERI-PEDES and TBAPR show improvements in Rank-1 accuracy and mAP over the CFAN baseline. The code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.