Joint Modality- and View-Aware Positional Residuals for Any-Scenario Person Re-Identification
Abstract
Any-Scenario Person Re-Identification (AS-ReID) aims to match individuals across diverse combinations of sensing modalities and camera viewpoints. Although existing transformer-based methods have achieved promising performance, they commonly share positional embeddings across sensing modalities. However, this shared positional prior does not explicitly account for modality-dependent spatial cues or the changes in body-part layout caused by aerial-ground viewpoint shifts. To address these limitations, we propose Modality- and View-Aware Positional Residuals (MVPR), a framework comprising two complementary components: Modality-Aware Positional Residual Correction (MPRC) and Text-Guided View Alignment (TGVA). MPRC learns modality-specific positional residuals and injects differences between consecutive layer-specific residuals at successive transformer blocks, enabling layer-wise, position-dependent corrections within a shared visual encoder. TGVA uses learnable text prompts to supervise paired forward passes with MPRC disabled and enabled. For aerial images, the two passes use aerial- and ground-view text anchors, respectively, whereas for ground images both use ground-view anchors. These asymmetric target assignments guide aerial-to-ground adaptation while maintaining ground-view supervision for ground images. Experiments on WHU-MARS-1000 demonstrate that MVPR surpasses UAD by 3.4% and 7.7% in mAP and Rank-1 accuracy, respectively, under the AS-ReID setting. The code will be available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.