Beyond Next-Token Prediction: Expanding Large Language Model Reasoning Coverage with Block Diffusion Objective
Abstract
Autoregressive (AR) language models have long been the dominant paradigm for modern large language models (LLMs). Recently, Diffusion Language Models (dLMs) have emerged as an alternative paradigm, distinguished by their ability to generate sequences in arbitrary orders. However, recent studies suggest that arbitrary-order generation may, in fact, limit the reasoning capabilities of dLMs relative to conventional autoregressive decoding strategies. This observation raises an intriguing possibility: if dLMs perform best under autoregressive decoding, can the diffusion training objective itself provide an alternative to the next-token prediction objective for training capable autoregressive models? Through a fully controlled experimental framework, we systematically investigate how the two training objectives shape reasoning capabilities across pre-training, mid-training, and RL post-training, evaluated on both extrapolative generalization and contextual generalization. Our experiments reveal several key insights into the impact of the diffusion objective. Firstly, the diffusion objective can yield stronger reasoning potential and enables greater gains during reinforcement-learning post-training. Secondly, diffusion-trained models exhibit stronger contextual generalization and greater robustness to reward hacking. Thirdly, the two objectives exhibit distinct advantages during mid-training: next-token prediction provides stronger gains on in-domain tasks, whereas block diffusion is more effective for out-of-domain generalization. Together, these findings highlight the potential of mask diffusion as an alternative training objective for reasoning language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.