Self-Regularized Flow Language Models
Abstract
We introduce SERF (Self-Regularized Flow Language Models), a simple self-supervised framework for improving continuous diffusion language models by exploiting the denoising trajectory itself. While continuous flow-based language models offer a promising alternative to autoregressive generation, their training objectives typically rely on supervision from clean target tokens, leaving the denoising trajectory underexploited as a learning signal. SERF addresses this limitation through noisy-to-clean self-distillation: an exponential moving average (EMA) teacher receives a cleaner version of the student's noisy input along the same interpolation path and provides additional supervision. We investigate two variants: SERF-L, which aligns the student's intermediate representations with deeper teacher representations, and SERF-P, which aligns their output token distributions. Both augment the standard training objective without requiring external pretrained models or modifying the inference procedure. Experiments across three continuous language-model paradigms (FLM, ELF, and S-FLM) demonstrate substantial improvements in training efficiency, generation quality, and few-step generation. In particular, SERF-L achieves up to 7.7× wall-clock training acceleration on OpenWebText, improves MAUVE by 24.1% and GSM8K accuracy by 26.0%, and matches baseline performance with much fewer sampling steps. These results establish noisy-to-clean self-distillation as an effective approach to improving continuous diffusion language models. Code and pretrained models will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.