How to Train Your Latent Flow Map for Few-Step Text Generation
Abstract
Diffusion language models generate tokens in parallel, but they need many denoising steps, which offsets their speed advantage. Flow maps reduce the number of steps by learning to jump between noise levels in a single network evaluation. Prior flow-map language models build these maps over token embeddings or frozen encoder states, and each makes its own design choices, so it is unclear which choices matter. We study them in a latent space learned jointly with the denoiser: we distill a latent diffusion language model into a flow map within its own latent space and analyze each component of the distillation, namely the regression target, self-conditioning, rollout depth, decoder alignment and the treatment of the latent space. Refining the target with the teacher matters most for open-ended text, while aligning the decoder to the flow map's outputs is essential on GSM8K, where a single wrong token breaks a program. The pre-trained latent space can stay frozen, and adapting it during distillation, together with the teacher, can improve coverage (pass@ at large ) but not single-sample accuracy, which drops on Sudoku. With the resulting recipe, distillation on GSM8K takes k iterations, and our model with steps exceeds its teacher with up to steps and matches the best baseline, which needs steps. On OpenWebText, with steps, it has % lower generative perplexity than the strongest flow-map baseline at matched entropy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.