acceptodds
Under review as a conference paper at ICLR 2027

From Global to Local: Rethinking CLIP Feature Aggregation for Person Re-Identification

Abstract

CLIP-based person re-identification (ReID) methods typically compress spatial evidence into a single global [CLS] representation, which can mix identity-relevant and corrupted regions under occlusion and cross-camera variation. We propose SAGA-ReID, a feature aggregation framework that delays spatial compression by reconstructing each patch through a shared anchor basis before pooling: image patches act as cross-attention queries while learned anchors serve as keys and values, preserving patch-indexed representations rather than first compressing the image into a small set of anchor-indexed features. This limits direct mixing between clean and corrupted patches during refinement and enables selective aggregation afterward. We evaluate this mechanism under synthetic masking and human-distractor occlusion, where SAGA's advantage over global pooling grows with occlusion severity before declining as identity information is exhausted. On standard and occluded ReID benchmarks, SAGA consistently improves over CLIP-ReID, with gains of up to +9.1 Rank-1 on occluded benchmarks. Its reconstructed feature achieves higher Rank-1 than CLIMB-ReID’s dedicated Mamba feature despite starting from a weaker backbone, while integrating SAGA with the CLIMB backbone further improves the final fused system on both Rank-1 and mAP. Finally, sensitivity analysis shows that performance is largely invariant to anchor-basis size, indicating that the gains arise from the reconstruction geometry rather than increased anchor capacity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.