Rethinking Distillation in Pedestrian Re-Identification: A Spatial Aggregation Perspective
Abstract
Pedestrian re-identification (Re-ID) is an open-set retrieval task in which practical systems often require either improving retrieval performance under fixed computational constraints or reducing inference cost with limited accuracy degradation. To address these requirements through knowledge distillation, we revisit which teacher signal should be transferred for Re-ID. We find that large changes in individual descriptor values do not necessarily imply comparable degradation of retrieval-relevant structure, motivating supervision of how descriptors are formed rather than their exact values. In Vision Transformer (ViT)-based Re-ID models, late-block [CLS]-to-patch attention provides an observable proxy for how spatial evidence is aggregated into the global representation. We study spatial aggregation as an effective teacher signal for open-set Re-ID, focusing on how the retrieval representation is formed rather than on exact agreement with teacher outputs. Across multiple ViT-based Re-ID frameworks and benchmarks, spatial aggregation transfer consistently improves both same-capacity refinement and lightweight student distillation, requiring only a teacher model and standard Re-ID training data, without additional annotations or inference-time components.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.