Understanding and Correcting Conformity in Intra-Modal Retrieval
Abstract
Frozen vision–language representations are widely used for intra-modal retrieval, yet their embedding geometry contains systematic shared structures that can distort cosine similarity and degrade ranking quality. We first show that conformity, the tendency of an embedding to be similar to many samples, is exactly determined by the embedding's projection onto a single shared direction defined by the normalized reference mean. We further show that mean correction alone is insufficient: after subtracting the reference mean, the -normalization implicitly imposed by cosine similarity in retrieval gives rise to a distinct residual conformity direction. Motivated by these observations, we propose Conformity-Guided Debiasing(CGD), a training-free geometric correction that attenuates conformity and residual conformity. CGD first applies mean correction and -normalization, and then performs a sample-dependent angular contraction toward the subspace orthogonal to the residual conformity direction. The resulting transform is plug-and-play, requires no backbone modification or query gallery fitting, and can be applied on top of existing intra-modal retrieval methods. Across multiple CLIP-like backbones and image-to-image and text-to-text benchmarks, our method consistently improves retrieval performance, showing that explicitly correcting conformity-related directional structure provides a simple and effective way to refine frozen intra-modal representations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.