Loopholing All the Way: Language Modeling with pure Self-Conditioning
Abstract
Recent work proposes continuous flow language models (FLMs) to address the limitations of discrete diffusion models (DDMs): at every step, DDMs sample tokens independently, which hurts few-step generation, and discard the uncertainty of the denoiser's predictions (*information collapse*). Empirically, on reasoning tasks such as Sudoku and GSM8K, FLMs at best match standard DDMs and fall well behind DDMs with self-conditioning (SC). In particular, Loopholing DDMs (LDDMs) mitigate information collapse by propagating an SC state across steps, along the discrete tokens. While SC is traditionally seen as an auxiliary heuristic to improve performance, its empirical success makes us wonder whether **we could define a continuous generative model based purely on self-conditioning?** We answer affirmatively with LAW (“Loopholing All the Way”), which propagates the *expected embedding* under the denoiser's predictions across steps, decoding the tokens once, at the end. On hard Sudoku, LAW comes close to LDDMs (90.4% vs. 92.7%), and on Maze, it is slightly more accurate than all baselines with few steps. On TinyGSM, fine-tuning a pre-trained Duo checkpoint with LAW increases the GSM8K accuracy from 18.6% to 55.2%. On OpenWebText, trained from scratch, LAW matches or improves over S-FLM with the standard DiT architecture.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.