Understanding and Mitigating State Exposure Mismatch in Flow Model Reward Alignment
Abstract
Reinforcement learning (RL) offers a flexible way to align pretrained flow models with task-specific objectives. Recent endpoint-based RL methods improve efficiency by reconstructing training states from rewarded endpoints, but these states may differ from policy-induced rollout states. We study this mismatch in state exposure during RL training of flow models. Our empirical analyses show that the discrepancy varies across timesteps, affects reward differently across denoising stages, and benefits from corrections aligned with individual trajectories. Guided by these observations, we propose Drift-Aligned Exposure Correction (DAEC), a simple method for mitigating this mismatch. DAEC perturbs only the trainable model input around analytically reconstructed states, using a readily available bridge residual to identify directions relevant to rollouts and adapting the perturbation magnitude to each timestep and sample, while preserving the original regression target and reference anchoring. DAEC introduces only lightweight tensor operations and requires no additional model evaluations per training update. Evaluations across diverse reward objectives demonstrate that DAEC improves both RL training efficiency and generation quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.