acceptodds
Under review as a conference paper at ICLR 2027

Text-Anchored Feature Deformation for Aerial–Ground Person Re-Identification

Abstract

Aerial-ground person re-identification requires matching identities across substantial changes in viewpoint and visual appearance. A central challenge is to establish cross-view correspondence while retaining the view-specific cues needed for identity discrimination. Current methods typically utilize direct feature alignment strategies, which do not explicitly model how visual evidence should be reorganized across viewpoints. Therefore, we propose Text-Anchored Feature Deformation (TAFD), which formulates cross-view matching through language-conditioned directional transformations in feature space. A Semantic View Anchoring (SVA) module learns aerial and ground semantic anchors using shared and view-private text prompts encoded by a frozen text encoder. Conditioned on the difference between these anchors, a Directional Feature Deformation (DFD) module employs two independent operators, one for aerial-to-ground transformation and the other for ground-to-aerial transformation, to predict spatial offsets, resample multi-level patch features, and aggregate target-view evidence. To preserve identity during transformation, a Target-Space Identity Consistency (TIC) objective supervises transformed descriptors using target-view classifiers and batch identity prototypes constructed from stop-gradient target features, while native features retain view-specific classification and metric supervision. Experiments on CARGO and AGReIDv1 demonstrate the effectiveness of TAFD for aerial-ground person retrieval. On CARGO, TAFD achieves 72.59% mAP and 76.60% Rank-1 for aerial-to-ground retrieval. On AGReIDv1, the same architecture achieves average bidirectional performance of 80.79% mAP and 87.51% Rank-1.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.