acceptodds
Under review as a conference paper at ICLR 2027

TE-VLM: Transfer Entropy-Inspired Distillation for Vision-Language Models

Abstract

Transfer entropy (TE) is a principled measure of directed information flow, but it is intractable to estimate in high-dimensional multimodal representation spaces. We use optimization-step TE, the information a frozen teacher provides about a student's next-step representations beyond the student's current state, as a conceptual lens for vision–language model (VLM) distillation and derive from it two tractable proxy objectives, TE1 and TE2. Both reward the student for reproducing the teacher's local embedding geometry by matching within-batch finite-difference directions up to orientation through squared cosine similarity, per modality (TE1) or jointly across image and text (TE2), with no density estimation, no extra teacher passes, and negligible compute. In CLIP-style distillation with RN50, ViT-B/16, and RN5016 teachers and RN34/RN18 students, the proxies consistently improve image-to-text retrieval over contrastive, KL, MSE, interactive-contrastive, mutual-information, and relational Gram-matching baselines while remaining competitive on text-to-image retrieval. With the base objective fixed, adding TE lifts MSCOCO image-to-text R@1 from 6.1% to 10.3%, against 8.7% for the strongest competing regularizer. The gains persist across temperatures, batch sizes, datasets (MSCOCO, Flickr8k, Flickr30k), and five random seeds; transfer better under distribution shift (MSCOCOFlickr8k); and extend to classification, where the distilled student surpasses its teacher in label-free Food-101 evaluation and generalizes better to ImageNet-1k zero-shot classification. Representation-level analyses show that TE-distilled students preserve the teacher's neighborhood structure and joint image-text geometry more faithfully than mutual-information distillation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.