FlowScope: Interpreting Generation Dynamics in Continuous Flow Language Models
Abstract
Continuous Flow Language Models (CFLMs) generate text by integrating a learned velocity field from noise to the sequences of embeddings that represent texts, providing an alternative to discrete diffusion and autoregressive language models. However, the generation dynamics in the continuous embedding space are not mechanistically studied. This work presents the first mechanistic interpretability analysis of CFLMs, examining the generative process through its central driving force, the velocity field, with sparse dictionary learning. Two embedding CFLMs, LangFlow and ELF, are investigated. We analyze the models based on their distinct designs, use the noise-free per-step estimation of the output state as the proxy for the model’s internal belief, and map the evolution of these beliefs along generation paths. Token identities commit progressively following the learned schedule, in a coarse-to-fine order across token types. The disentangled dictionary separates a small population of frequently activating features from specialized features whose activity increases as trajectories converge, and activation steering then verifies the mechanical causality in these flow models. By examining per-variant trajectory properties, we implement online stability stopping for LangFlow and static schedule truncation for ELF, reducing denoiser evaluations by 12–25% while achieving 0.86–0.94 paired-token agreement, with measurable quality trade-offs. The analysis further examines the dynamics of self-distilled ELF weights, noting observable shifts in their commit clock and feature activation profile.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.