Length from the Latents: Dynamic-Length Compressed Latent Diffusion Models
Abstract
Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive models, offering parallel generation and the ability to revise their output through bidirectional attention. However, the adoption of DLMs is still limited due to running diffusion on high-dimensional token-level representations. Some methods overcome this issue by running diffusion in a compressed fixed-dimensional latent space, where the latent space is realized with a compressor network, and later a decompressor network expanding the denoised latents back to token-level representation followed by decoding them to the output tokens. Yet, current methods run parallel decompression assuming a fixed output sequence length, and thus attend bidirectionally over all output positions, including the ones that are beyond the end-of-sequence token. This results in every request paying for the maximum length, and in compute wasted on padding. As a solution, we propose PACED (Parallel Append-only Cached Extensible Decompression). PACED masks the decompressor so each output position attends to all latents but only to itself and to preceding positions, ensuring prefix-consistent decompression that is efficiently extendable. We realize this with the help of training a length head on the denoised compressed latents, the features that are natively length-aware. Post token decoding, if the end-of-sequence token does not appear, we reuse cached internal states of the decompressor, similar to key-value caching, and decompress only the missing positions, instead of entirely rerunning the decompression. We prove this returns exactly the text a full-length decode would have returned, whatever initial length was chosen. Distilling the pretrained decompressor into the prefix-consistent one on sampled latents recovers much of the generation quality while allowing PACED to run batched 1.2–1.8 faster than full-length decompression. PACED thus makes decompression cost follow the output rather than the maximum length, paving the way for compressed-latent DLMs that support long outputs efficiently.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.