PANDA: Training-Time Privileged Attribute Supervision for Ship Re-Identification
Abstract
Ship Re-Identification requires distinguishing similar vessels across cameras. Identity labels tell the model which images match, but not which visual details matter. Language can provide this guidance, yet sentence-level alignment does not explicitly teach attribute discrimination, while MLLM-based reranking adds inference cost. To address these problems, we propose **P**rivileged **A**ttributes for **N**ative **D**escriptor **A**daptation (**PANDA**), a three-module framework that only uses language supervision during training. Specifically, **P**rivileged **A**ttribute **C**onstruction (**PAC**) extracts shared attributes from paired training views, organizes them into structured descriptions, and creates candidate negatives by replacing one slot with text from another identity. Using these annotations, **F**rozen-LM **A**ttribute **S**upervision (**FAS**) trains visual adapters to generate descriptions and distinguish matching attributes from alternatives through a frozen language model. **A**ttribute-**A**ware **F**usion (**AAF**) then combines identity and semantic features from the attribute-supervised encoder into a single retrieval descriptor. At inference, only the visual encoder and fusion head remain, requiring neither text nor a language model. Experiments on ShipReID-2400 and VesselReID show that PANDA outperforms the compared baselines while maintaining efficient inference, and ablation studies demonstrate the effectiveness of its components.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.