SimpLM: Simplex-Based Flow Language Modeling via Autoregressive Adaptation
Abstract
Autoregressive language models remain the dominant language modeling paradigm, yet their sequential decoding limits efficiency. While recent efforts have explored discrete diffusion language models for parallel decoding, these approaches suffer from severe degradation in text quality when decoding multiple tokens simultaneously due to token-wise independence assumptions. Continuous-space diffusion and flow models offer a compelling alternative, but existing methods are often restricted to unconditional full-sequence generation, and some rely heavily on learned embedding encoders to map discrete tokens into continuous latent states, requiring additional pretrained encoders. In this work, we propose SimpLM, a continuous flow language modeling framework that operates directly on the probability simplex in the vocabulary space. By applying Gaussian interpolation to the probability states, we derive a closed-form posterior logit correction that yields an analytic skip connection for logit transfer, bypassing the hidden-state information bottleneck in Transformer backbones. Beyond this, SimpLM natively supports conditional and block-based semi-autoregressive generation via a vectorized time formulation that assigns separate time variables to block positions, providing a principled foundation for continuous-time conditional generation. Furthermore, SimpLM can adapt a pretrained autoregressive model into a continuous flow model, reducing the cost of training non-autoregressive models from scratch. Empirically, with only 30B training tokens, SimpLM achieves a generative perplexity below 20 at block size and generation length of 128.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.