GapDIM: Dimension Pruning for Modality Gap Reduction in Vision-Language Models
Abstract
Contrastive Vision-Language Models (VLMs) such as CLIP and SigLIP exhibit a modality gap, whereby embeddings from different modalities occupy separate regions of the shared latent space. Although this fragmentation does not necessarily prevent strong retrieval performance, it limits the use of pretrained embeddings for out-of-distribution or group-wise tasks such as clustering. Existing methods to reduce the modality gap typically require retraining with additional losses or architectural changes, while post-hoc alternatives often degrade retrieval by disrupting the instance-wise alignment learned during pretraining. In this paper, we ask whether the portion of the gap that disrupts the multimodal organization of the representation space can be removed selectively. We observe that the gap is concentrated in a few coordinates of the native bases of pretrained models and that these coordinates exhibit low participation in the dominant centered subspaces. We first separate these two geometric phenomena and relate them through a gap-leverage bound. Based on this analysis, we propose Gap-Importance DIMension pruning (GapDIM), a post-hoc, training-free method that identifies coordinates with high modality gap and low subspace participation and selectively prunes them from the embedding space. GapDIM thereby reduces the modality gap while preserving the discriminative structure of pretrained representations. Across multiple VLM backbones and datasets, GapDIM generally preserves retrieval performance while reducing modality separation and improving joint clustering in most configurations. These results suggest that the largest part of the modality gap is concentrated in a small set of low-participation embed- ding coordinates, enabling a training-free calibration strategy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.