acceptodds
Under review as a conference paper at ICLR 2027

Cross-Modal Feature Dynamics Distillation for Resource-Constrained Multimodal LLMs

Abstract

Multimodal large language models (MLLMs) achieve strong performance on vision-language tasks, but their substantial computational cost limits deployment in resource-constrained settings. Knowledge distillation provides a promising route for transferring cross-modal reasoning knowledge from a large teacher to a compact student. Existing MLLM distillation methods, however, predominantly match static outputs or representations at isolated layers, overlooking the layer-wise evolutionary nature through which cross-modal reasoning is progressively formed. We introduce Cross-Modal Feature Dynamics Distillation (CM-FDD), which distills cross-modal grounding from two complementary perspectives: the grounding state at each depth and its evolution across selected depths. To enable this transfer across heterogeneous teacher-student models, we represent grounding with text-to-vision affinity, providing a model-comparable space that captures both where visual evidence is grounded and how that grounding changes over depth. When teacher and student use different visual resolutions, we further align their visual supports through deterministic cross-resolution projection. Experiments on ten benchmarks show that CM-FDD consistently outperforms prior MLLM distillation methods and achieves state-of-the-art performance under matched model-size settings, demonstrating the effectiveness of preserving grounding dynamics for efficient multimodal models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.