Latent Flow Transformer: Flow Matching the Hidden Dynamics of Language Models
Abstract
Transformers, the standard implementation for large language models (LLMs), typically consist of tens to hundreds of discrete layers. While depth improves performance, it is an inefficient use of capacity next to the continuous formulations that diffusion and flow-based models use for image generation. We propose the Latent Flow Transformer (LFT), which replaces a contiguous block of layers with a single transport operator trained by flow matching, a post-training procedure that leaves the surrounding architecture unchanged. To identify replaceable spans, we introduce the Recoupling Ratio, an optimal-transport measure of how far a span's endpoint latents deviate from an optimal pairing. We further introduce Flow Walking (FW), which addresses the failure of existing flow-based methods to preserve latent pair coupling. On Pythia-1.4B, an LFT trained with FW replaces 12 of 24 layers with a single operator and retains 92% of the teacher's downstream accuracy (57.5% vs. 62.8%) at 63% of its parameters and 72% of its per-token FLOPs, whereas skipping the same span attains 53.8%. FW matches a regression transport layer at six replaced layers, surpasses it at twelve, and is less sensitive than standard flow matching to the latent coupling of the replaced span, bridging the gap between autoregressive and flow-based generation paradigms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.