OVERTURE: Decode Before the Echo for dMLLM
Abstract
Existing diffusion-based multimodal language models (dMLLMs) typically order decoding by prediction confidence. This strategy prioritizes high-confidence tokens, providing more reliable context for later predictions, but overlooks how decoding timing affects multimodal decision reliability. Analyzing complete denoising trajectories, we find that confidence-based decoding finalizes simple, linguistically predictable tokens first, while deferring more difficult tokens whose interpretation requires precise cross-modal alignment and richer semantic descriptions than those conveyed by formulaic, template-like language. As finalized text accumulates, linguistic context increasingly shapes later predictions, creating a growing linguistic echo, whereas visual evidence exerts less influence on candidate competition. For these visually dependent and more difficult tokens, reliable decisions require stronger visual evidence. Yet confidence-based decoding postpones their commitment until visual evidence exerts weaker influence, creating a difficulty–timing mismatch between evidence requirements and decision timing. This mismatch makes visually dependent decisions more vulnerable to erroneous final predictions. Tracking each token's attention context across denoising steps reveals a stable local Undershoot phenomenon. Pronounced Undershoot is concentrated among tokens associated with subsequent erroneous decisions, providing an observable signal for locating vulnerable generation decisions. We propose OVERTURE, a training-free, trajectory-guided priority decoding method that identifies vulnerable content from Undershoot and prioritizes the corresponding positions during regeneration while re-predicting tokens from the current generation state. Experiments across multiple benchmarks and model backbones demonstrate the broad effectiveness of OVERTURE.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.