Into the Continuum: Efficient Semi-Continuous Diffusion LLM
Abstract
Continuous diffusion language models retain partially resolved token information across denoising steps. We study whether pretrained autoregressive (AR) LLMs can acquire this generation mechanism with limited additional data and training. Our efficiency focus is this conversion budget; cross-model inference speed remains to be measured. We present Semi-Continuous Diffusion Language Model (SC-dLLM), the first low-data conversion of pretrained AR LLMs to blockwise continuous-state generation. Each block starts from masks, evolves through continuous states, and yields discrete tokens. A brief continuous endpoint initialization prepares the backbone for same-position prediction. The main adaptation then trains on independently corrupted intermediate states while retaining the vocabulary prediction objective. During generation, token predictions guide stochastic updates and determine when confident positions are committed. Training uses analytically sampled states without an online teacher or unrolled generation trajectories. The 7B conversion uses approximately 1.1B additional training tokens, about of the 45.2B effective tokens reported for ELF's OpenWebText training. Across code generation, mathematics, instruction following, and knowledge-intensive QA, SC-dLLM reaches nine-score averages of 45.2 at 1.5B and 61.2 at 7B, exceeding the reported AR-initializer and Fast-dLLM v2 averages at both scales. At 7B, all four HumanEval and MBPP Base/Plus scores exceed both the same-data AR fine-tuning control and Fast-dLLM v2. These results establish the feasibility of adapting AR backbones to blockwise continuous generation with a small additional training budget using the pretrained vocabulary head.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.