Alignment-Preserving Low-Dimensional Projection for Efficient Vision-Language Retrieval
Abstract
Vision-language models produce high-dimensional embeddings that enable effec- tive cross-modal retrieval but incur substantial storage and similarity search costs at scale. Naive post-hoc dimensionality reduction alleviates these costs but may substantially distort the cross-modal geometry required for retrieval. We introduce Alignment-Preserving Low-Dimensional Projection (ALP), a lightweight post-hoc framework that maps frozen vision-language representations into a compact shared embedding space using modality-specific projections. We further introduce ALP+KD, which distills the cross-modal similarity distribution of the original representation into the compressed space. On our primary CLIP retrieval benchmark, ALP+KD compresses 512- dimensional embeddings to 128 dimensions while improving Mean Recall@1 from 0.4389 ±0.0025 to 0.4727 ±0.0046 across three random seeds. At the same dimensionality, PCA achieves 0.3295 ±0.0032 and a parameter-matched Matryoshka-style baseline achieves 0.4543 ±0.0035. The compact embeddings also provide practical systems benefits. On a 500K- vector FAISS database, ALP+KD reduces raw vector storage by 75%, accel- erates exact inner-product search by approximately 2.9×, and reduces index- construction time by 4.02×. Zero-shot cross-dataset and cross-backbone exper- iments further indicate that most of the original retrieval structure is retained outside the primary training setting. These results show that explicitly preserv- ing cross-modal similarity geometry can provide a favorable accuracy–efficiency trade-off for multimodal retrieval
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.