VisWave: Revisable Wavefront Decoding for Diffusion Vision-Language Models
Abstract
Diffusion multimodal large language models (Diffusion MLLMs) leverage bidirectional context modeling and intra-block multi-token denoising to provide an efficient text-generation path distinct from autoregressive decoding. However, long-form generation remains constrained by serial dependencies across blocks, causing inference efficiency to degrade as the output grows. Existing acceleration methods mainly reduce intra-block denoising cost and do not exploit computation that can overlap between adjacent blocks. The key obstacle is that an early-started successor depends on an unstable predecessor context and may become invalid as that context is updated. We find that this effect is usually local: most successor draft tokens remain reusable, while only a small number require re-verification.Based on this observation, we introduce VisWave, a training-free decoder for diffusion vision-language models. VisWave overlaps the decoding of adjacent blocks through revisable wavefront execution and uses virtual-remask token-to-token verification (VR-T2T) to re-evaluate affected draft tokens under the latest textual and visual context. Supported tokens are retained, while the remaining positions are remasked and returned to the native denoising process. This local selective revision makes inter-block parallelism reliable. Across seven natural workloads, VisWave accelerates BARD-VL-4B by , with only a 0.27 percentage-point decrease in macro accuracy across six accuracy tasks; at a fixed generation length of 192 tokens, it reaches speedup. The efficiency gains remain stable from 2B to 8B and transfer to SDAR-VL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.