Large Flow Language Models
Abstract
Flow and diffusion models are leading approaches to generative modeling of continuous signals such as images, videos, audio, and protein structures. Their parallel generation makes them attractive for fast language modeling, but it conflicts with the inherent directionality of text where earlier tokens contextualize the later ones, whereas conventional flows denoise all positions synchronously. As a consequence, flow language models have largely remained at sub-billion scale and have not established competitive quality-speed tradeoffs on realistic tasks. We present Lafa, a block-autoregressive flow language model that preserves causal order across blocks while learning an asynchronous progression within each block, and scale it to 8B parameters by converting a pretrained autoregressive model using less than 10B training tokens. Lafa retains both a block denoiser and a next-token predictor, unifying sequential next-token prediction, parallel flow decoding, confidence decoding, and self-speculative decoding, in a single model. Across six benchmarks spanning instruction following, reasoning, and code, these modes trace a broad quality-speed Pareto frontier, and Lafa attains 2.54x the throughput of a strong autoregressive baseline at matched quality. Our findings show that continuous flows can be a practical and scalable foundation for language modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.