acceptodds
Under review as a conference paper at ICLR 2027

Continuous Space Drifting for Discrete Diffusion Language Models

Abstract

Masked diffusion language models have become capable generators, motivating objectives that complement token reconstruction with sequence-level feedback. We present LD2LM (Large Drifting Diffusion Language Model), which adapts continuous feature-space supervision to an 8B discrete block-diffusion model. Building on TokenDrift's differentiable soft-token interface, LD2LM retains denoising cross-entropy (CE) and adds multiscale feature views across encoder depths and sequence regions. We also study how this supervision interacts with token- and feature-level KL constraints. Although TokenDrift reports benefits from feature-space Drifting, directly transferring its Drift-only recipe to DreamReasoner8B performs substantially worse than CE. We hypothesize that Drift-only feedback may not transfer directly to scaled discrete block-diffusion models. Drawing on scaling experience with richer representations, LD2LM combines continuous multiscale feature feedback with token-level CE. In a matched training-stage comparison, the Full schedule raises HumanEval from 74.39% to 75.00%, MBPP from 60.40% to 61.20%, and LiveCodeBench pass@1 from 44.84% to 49.26%; tokens per forward pass increase from 1.5490 to 1.5679, with the inference sampler unchanged. In a separate token-KL comparison, adding Drifting improves HumanEval from 73.78% to 75.61% and LiveCodeBench from 49.56% to 51.33%, while tying MBPP at 61.40%. These results show the promise of continuous multiscale feature feedback for adapting Drifting to scaled discrete block-diffusion models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.