S²I-dLLM: From Sparsity and Specialization to Adaptive Acceleration in Multimodal Diffusion LLMs
Abstract
Multimodal diffusion large language models (dLLMs) integrate tasks across modalities within a shared bidirectional Transformer and a masked diffusion objective, yet how the computational demands of different tasks are organized within the shared network and evolve during denoising remains insufficiently understood. In this work, we analyze the computational demands of multimodal dLLMs from two perspectives, network structure and token representation evolution, characterizing how different tasks sparsely depend on shared computation and how tokens with different roles differ in representational stability and update requirements. Building on this analysis, we propose S²I-dLLM, a training-free, task-aware acceleration framework that adaptively allocates computation according to the demands of different tasks and how these demands change throughout denoising. S²I-dLLM reduces network channels and selects layers to execute according to task requirements, while restoring necessary computation as denoising progresses to control the accumulation of approximation errors. At the token level, the framework uses representational stability to guide feature reuse and prediction states to adjust the timing of output refinement and commitment, thereby coordinating per-step computational cost with generation progress. Experimental results demonstrate that S²I-dLLM effectively improves the inference efficiency of multimodal dLLMs while preserving task performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.