The Generative Gap Between Speculative Drafter and Continuous Diffusion Language Model
Abstract
Continuous diffusion language models (CDLMs) and speculative decoding offer two approaches to efficient language generation. Yet whether a strong drafter can also serve as a competitive standalone generator remains unclear. We instantiate CDLMs with LangFlow and speculative decoding with DSpark, systematically comparing LangFlow with standalone adaptations of DSpark’s parallel-backbone architecture on unconditional OpenWebText generation at lengths 16, 128, and 1024. We evaluate both paradigms using two complementary metrics—generation perplexity (GenPPL) and MAUVE—to assess evaluator predictability and distribu- tional similarity. A conditional KL analysis identifies dependencies inaccessible to DSpark’s fixed-backbone Markov head. At lengths 128 and 1024, DSpark-Markov trails LangFlow on both metrics. To test whether richer access to sampled history can narrow this gap, we evaluate DSpark-RNN, a recurrent-head extension of the Markov drafter. Although recurrence improves both metrics at every tested length, DSpark-RNN remains behind LangFlow at lengths 128 and 1024; at length 16, RNN instead leads. Teacher L1 supervision lowers Markov GenPPL without im- proving MAUVE, while additional denoising steps improve LangFlow’s scores at fixed weights. These results indicate that richer history access narrows but does not close the observed longer-sequence gap: learning target-compatible token propos- als under observed prefixes remains distinct from modeling a coherent distribution over future tokens in free-running generation. Our findings clarify the distinction between drafting efficiency and generative competence, and highlight the strength of LangFlow’s iterative continuous modeling for high-quality, verifier-free parallel language generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.