OSFAL: Orthogonal Cross-Spectral Feature Alignment and Layer-Localized Cross-View Learning for Any-Scenario Person Re-Identification
Abstract
In Any-Scenario Person Re-Identification (AS-ReID), a unified model must retrieve identities across spectral modalities and aerial–ground views. However, existing methods handle these variations through joint scenario alignment without explicitly accounting for their distinct characteristics, potentially limiting retrieval performance. To address this limitation, we propose OSFAL, which combines modality-specific orthogonal alignment with layer-localized cross-view learning to mitigate spectral and view discrepancies through distinct mechanisms. Specifically, Orthogonal Cross-Spectral Feature Alignment (OCSFA) combines orthogonal rotations with mean alignment to map features from different spectral modalities into a shared reference coordinate system while preserving within-modality Euclidean distances. We learn these transformations by encouraging the modality–view prototypes of each identity to agree with a identity consensus prototype. Layer-Localized Cross-View Learning (LLCVL) compares within-view and cross-view similarities across encoder blocks, together with cross-view identity matching accuracy, to locate a key supervision boundary. Bidirectional cross-view contrastive learning is then applied at the selected block output within each modality. Experiments on WHU-MARS show that OSFAL improves retrieval performance across different evaluation protocols, achieving +5.0% gain in Rank-1 accuracy over the baseline on WHU-MARS-1000.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.