acceptodds
Under review as a conference paper at ICLR 2027

Scalable Information Flow Distillation for Large Multimodal Language Models

Abstract

The application of Multimodal Large Language Models (MLLMs) in resource-constrained scenarios is significantly limited due to their enormous parameter size and computational complexity. Knowledge distillation aims to transfer knowledge from large models to smaller ones, striking a balance between performance and capacity. Recent works have gradually shifted their focus from response-based distillation to intermediate-layer-based distillation, with the core objective being to explore the intrinsic working mechanisms of MLLMs. In this paper, we propose a new distillation paradigm based on the information flow of MLLMs. Specifically, through systematic analysis, we reveal the information flow pathways between text, images and the last token (used for generation) as they pass through the intermediate layers. Consequently, we identify a three-stage workflow for MLLMs: understanding the text firstly, then processing the multimodal data and finally focusing on generating a response. Moreover, this information flow exhibits a proportional scaling property with respect to the number of layers across models of different sizes. Based on these findings, we propose a simple distillation scheme that compresses the information flow of large models into small models. Extensive experiments verify the effectiveness of the proposed method.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.