acceptodds
Under review as a conference paper at ICLR 2027

The Denoising Estimate Is a State Variable: Joint Refinement for Few-Step Flow Language Models

Abstract

Flow language models (FLMs) generate text by repeatedly predicting a clean representation with a denoiser and using it to transport a noisy latent state toward the data distribution. The first-order samplers commonly used in FLMs evaluate the denoiser once at the current state and keep the resulting estimate fixed throughout the update. This approximation becomes inaccurate for large steps, because changing the estimate also changes the transported trajectory and hence the state at which subsequent predictions are evaluated. We introduce , which refines the estimate jointly with the state it induces. Given a candidate estimate, AIM-FLM constructs the induced intermediate state through the closed-form affine transport of the flow, re-evaluates the denoiser at that state, and repeats this Picard iteration toward a joint fixed point. Under linear interpolation schedules this fixed point corresponds to an implicit-midpoint step, but it is solved directly in the denoising-estimate space, preserving a clean representation that can be decoded at the endpoint and retained as feedback across solver steps or distilled hops. Under a local contraction condition, the iteration converges geometrically, and under a margin condition, the decoded tokens stabilize after finitely many iterations without requiring exact convergence of the continuous estimate. The same coupling applies to pretrained deterministic and stochastic samplers without retraining, can be preserved in flow-map distillation by passing the estimate from one step to the next, and can be incorporated into joint training of the denoising field and hop model. Across several flow language model backbones, AIM-FLM improves the trade-off between generation quality and compute. Without retraining, AIM-FLM surpasses the 32-NFE ELF-B sampler in generative perplexity using 17 denoiser evaluations, with comparable entropy and lower repetition. With joint training, it surpasses the same baseline using only 12.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.