CRED: Cross-Layer Participation Regularization Against Token Collapse in Vision-Language Embedding Distillation
Abstract
Vision-language embedding models map multimodal inputs into a shared semantic representation space. Embedding distillation transfers multimodal representation knowledge from larger teacher models to more compact students, but can lead to token representation collapse in the student. We find that the effective rank of image-token representations in the student undergoes a nearly order-of-magnitude reduction relative to its initial hidden state, whereas the teacher exhibits substantially milder changes across its depth. To address this discrepancy without requiring intermediate teacher activations, we introduce CRED, a teacher-independent regularizer based on the participation ratio (PR). CRED constrains the PR of selected student hidden layers to remain no lower than that of the initial hidden state, thereby preserving effective dimensionality throughout the network. Unlike layer-wise feature or covariance matching, CRED operates on a scalar spectral statistic that is invariant to representation rotations and requires no additional teacher inference. Moreover, PR can be efficiently approximated using random projections, avoiding eigendecomposition and scaling more favorably to inputs with large token counts. Combined with output-level distillation, CRED reduces layer-wise token collapse and improves performance on downstream multimodal tasks in MMEB over competitive baseline methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.