PACE: Prune Aggressively and Commit Early for Diffusion VLM Inference Acceleration
Abstract
Diffusion vision-language models (DVLMs) have emerged as an alternative to autoregressive VLMs for their capacity for parallel decoding. In practice, however, each denoising step runs full bidirectional attention over the entire sequence, and the resulting inference cost scales as the product of the step count and the per-step cost of a full-sequence pass. Since visual tokens make up most of the sequence, compressing them has become an active direction. We show that the step count is a design factor as important as compression for DVLM inference acceleration, and that the two are not separable. Compression is applied after the first denoising step, which commits part of the answer on the full visual context, and the step count determines the size of that part. Under the default schedule, the core content of the answer is determined on the pruned context, where aggressive compression loses its accuracy. With fewer steps, more of the answer is committed before pruning and less of it is exposed to the loss from compression, so penalty shrinks. We formalize this regularity as a monotone comparative statics result. Based on these observations, we propose PACE, a training-free inference acceleration framework for DVLMs that scores visual tokens from the first forward pass and selects the step count jointly with the retention ratio under a task-level rule. Reducing the step count or pruning alone degrades accuracy, whereas combining them removes most of the penalty, and the temporal lever improves existing DVLM token pruning methods as well. PACE achieves 97.05% mean relative performance across nine benchmarks, while reducing generation FLOPs to 13.6% of vanilla at the 10% token budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.