acceptodds
Under review as a conference paper at ICLR 2027

When Anchors Commit, Coherence Breaks

Abstract

Diffusion multimodal large language models (dMLLMs) generate responses through iterative denoising, with predictive confidence determining which masked positions are committed first. Under full-sequence confidence-based decoding, we identify a three-phase commitment pattern in the evaluated trajectories: opening sentences first, tail anchors next, and middle content last. This timing imbalance leaves extended masked spans in the middle of a response whose surrounding content is already fixed, and is associated with semantic discontinuities. We investigate local attention concentration as an intervention target: sequence-wide detection can prioritize dominant peaks at already committed positions while overlooking concentrations in unresolved regions. We propose GADS (Gap-Aware Dynamic Suppression), a training-free intervention that detects concentrated attention columns within a local window and applies a reference-preserving logit bias, while retaining full-sequence token selection. GADS shifts commitment toward previously delayed middle content and reduces the tail-before-middle gap from 31.6 to 6.7 global denoising steps on LLaVA-Bench. Across LLaDA-V, MMaDA, and LaViDa, it improves LLaVA-Bench scores by 12.0, 21.3, and 2.1 points, respectively, while reducing object hallucination rates on all three backbones. These results support local attention intervention as an effective approach to improving commitment trajectories and long-form multimodal generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.