acceptodds
Under review as a conference paper at ICLR 2027

Little Context, Large Guidance during Early Sampling in Masked Diffusion Language Models

Abstract

Masked diffusion language models apply classifier-free guidance (CFG) using a null prompt prediction as the reference, even during initial sampling when little response context is available. We show that this period forms a distinct guidance regime. The null prompt prediction changes rapidly, its alignment with the conditional prediction is weak, and the relative CFG residual strength is elevated. An algebraic decomposition and hidden representation analysis relate this elevated strength to differences between the two branches. We evaluate two adaptive controls for this regime. Adaptive on/off removes the early residual, whereas AdaptiveCap retains its direction while limiting its strength using a prompt specific reference. Both return to standard CFG once branch similarity reaches a threshold. Across two models, three benchmarks, and five guidance scales, AdaptiveCap consistently improves or matches Full CFG, with gains of up to percentage points. Comparisons with image generators show that elevated residual strength alone does not predict gains from early control. Our results indicate that when strong residuals coincide with weak branch alignment, retaining early guidance at a moderated strength is more effective than applying it unchanged or removing it entirely.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.