Manta-LM: Parallel Text Generation with Locally Structured Latents
Abstract
Parallel text generators update many positions at once, yet they often must respect text that is already fixed, such as a prompt, a source sentence, or the ending of a passage. Manta-LM generates in a continuous latent space whose positions stay tied to regions of the text. A convolutional TextVAE encodes tokens with local receptive fields: at 4× compression it reconstructs 99.75% of tokens on average across 19 datasets, and joining the latents of two separately encoded texts adds 0.45 percentage points of token error. A Transformer flow, trained by flow matching, keeps the latents of observed text fixed and jointly updates the free ones, which are then decoded in parallel; changing the observation mask switches between generation, continuation, and infilling. With one 101M flow per task, Manta-LM obtains the highest BLEU among the compared models on paraphrasing, question generation, simplification, and dialogue. Choosing classifier-free guidance for each step budget lets four sampling steps match the BLEU of 64 steps with the default guidance on question generation. In unconditional generation, 16 steps reach Gen PPL 46.9 versus 78.3 for RADD at similar unigram entropy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.