Contrastive Geometry Distillation for Vision–Language Dual Encoders
Abstract
We propose contrastive geometry distillation (CGD) for vision–language dual encoders. A decomposition of the teacher–student mutual information gap motivates two variants that combine feature distillation with KL matching of caption distributions for each image. CGD-N also matches distributions over unmatched images for each caption, while CGD-IW corrects selected unmatched pairs using teacher-to-student probability ratios. For CGD-N, exact matching recovers the teacher’s similarity matrix up to a global additive constant. Experiments show that CGD lowers similarities for negative pairs more than for positive pairs, widening retrieval margins while retaining alignment of image features with the teacher. The added losses supply small gradient corrections in directions distinct from feature distillation. The resulting students preserve the teacher’s candidate ordering more closely than CLIP-KD and CLIP-RD, with improved average retrieval and comparable classification accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.