Tilted Generator Matching: One RL Objective For Every Diffusion and Flow Model
Abstract
Diffusion and flow models are a flexible framework for generative modeling: they can corrupt data in almost any way, live on almost any state space, and generate with almost any sampler. This flexibility, however, makes them hard to post-train. Reinforcement learning (RL) methods for diffusions and flows have been built one at a time, and none applies across this zoo of models. To fix this, we introduce Tilted Generator Matching (TGM), a single RL objective for every model expressible in the generator matching framework. This encapsulates nearly all standard generative models, such as flow matching, score-based diffusion, and discrete diffusion. Our key insight is that RL reduces to weighted pretraining: each step of KL-regularized RL is solved exactly by the model's own pretraining loss, tilted by exponentiated reward. TGM is therefore likelihood-free, works with any sampler, and on-policy. We study TGM on diffusion and flow language models (DLMs), which span discrete and continuous state spaces and masked, uniform, and Gaussian noise. TGM is the first RL method demonstrated on uniform-noise DLMs and the first designed for Gaussian DLMs, as well as for few-step flow-map language models. On Sudoku, GSM8K, and sentiment tasks, a single configuration improves every DLM family and matches or beats family-specific methods, in one case more than doubling accuracy ( on GSM8K for uniform DLMs). On flow maps, TGM-Map lifts the entire accuracy–NFE frontier, matching base accuracy with up to fewer sampling steps, and learns from rollouts that are up to cheaper to generate, at no cost in final accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.