acceptodds
Under review as a conference paper at ICLR 2027

Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models

Abstract

Masked diffusion language models (dLLMs) have shown strong potential for faster inference through parallel token generation when combined with confidence-based samplers. However, recent work has shown that such methods can defer unmasking high-entropy *fork* positions at which multiple plausible continuations exist. This results in reduced generation diversity, as shown by worse pass@ scaling, and limits gains obtainable from RL post-training. To avoid this *flexibility trap*, prior work advocated for autoregressive (AR) sampling. Here, we show that discarding confidence-based sampling is unnecessary and, once inference cost is taken into account, wasteful. We first propose *Fork-dLLM*, a simple hybrid sampler that uses AR-style ordering only at uncertain fallback steps while retaining parallel generation otherwise. We then extend the same principle to post-training with *ForkGRPO*, which uses Fork-dLLM rollouts and applies the GRPO objective only at fallback steps, preserving exact policy-likelihood ratios while substantially reducing rollout and optimization cost. In our experiments, Fork-dLLM matches the strong pass@ scaling of AR sampling while being – more efficient, and ForkGRPO achieves downstream performance comparable to or better than AR-based GRPO baselines at a substantially lower training cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.