Directional Phase Transitions in Reward Guidance for Flow models
Abstract
Recent advances in reward-guided generation with flow models have demonstrated remarkable performance in steering outputs toward desired reward functions. In particular, inference-time alignment methods are drawing attention, as they adjust the sampling trajectory without extensive retraining of the pretrained model. Recent analyses of sampling trajectories, however, reveal that they are not dynamically homogeneous but pass through distinct phases. Yet, the impact of these phases on inference-time reward alignment remains under-explored. To bridge this gap, we introduce the directional expansion rate, a metric that quantifies how the velocity field of a flow model amplifies or attenuates perturbations toward higher reward. We observe a consistent contraction–expansion–contraction pattern across distinct tasks: data assimilation, PDE solving. To explain this behavior, we derive the phase structure in a closed-form expression for a Gaussian mixture model, establishing a formal link to mode commitment. Based on this analysis, we propose an adjoint-based control strategy that operates exclusively within the expansion window to steer mode selection before the commitment occurs. Across the tasks, we demonstrate that the proposed method generates better-aligned samples with the target reward function.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.