acceptodds
Under review as a conference paper at ICLR 2027

AdaVTP: Prediction Guided Adaptive Visual Token Pruning for Efficient Diffusion MLLM Inference

Abstract

Diffusion multimodal large language models (dMLLMs) enable strong multimodal understanding through iterative parallel decoding and bidirectional context modeling, but repeatedly process all visual tokens at every denoising step, causing substantial redundant computation. Existing methods mitigate this cost by pruning visual tokens using relevance estimated from early prediction steps, but apply a fixed retention ratio to every input. This overlooks varying visual evidence requirements: simple questions may require only a few local cues, whereas complex ones require information distributed across multiple regions. To address this, we propose AdaVTP, a training-free framework for adaptive visual token pruning in dMLLMs. After the first denoising step, AdaVTP uses prediction confidence to adaptively balance visual attention from the question and masked response positions, yielding task-conditioned relevance scores. Guided by these scores, AdaVTP first retains a sample-specific candidate set covering a prescribed fraction of the cumulative relevance, and then progressively selects complementary tokens by relevance-weighted semantic coverage gain until the target coverage is reached. The resulting subset is reused for all remaining denoising steps, avoiding repeated pruning overhead. Experiments on three dMLLMs and 13 image/video benchmarks show consistent gains over existing training-free pruning methods under matched dataset level token budgets. With only 10% of visual tokens retained overall, AdaVTP preserves 92.1% of the full token performance on LLaDA-V, exceeding RedVTP by 4.6 points. At comparable task performance, it delivers 1.39–1.92 end-to-end inference speedups over RedVTP.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.