Exploring the Design Space of Reward Backpropagation for Flow Matching
Abstract
Direct reward backpropagation aligns text-to-image flow matching models sample-efficiently, but practical methods fix different surrogate backward trajectories, leaving unclear which choices drive performance. Full-rollout differentiation requires substantial activation memory and can amplify gradients through chained Jacobians. Surrogates reduce these costs but introduce fidelity issues: detached methods query rewards on model-implied clean estimates, while connector methods such as LeapAlign use derivatives of single-velocity interval approximations. We propose FlowBP, a framework that treats the backward trajectory as the design object and exposes four axes: reward-model input, active set, integration weights, and bridge coupling. Prior direct-gradient methods occupy specific settings in this space. Transplanting endpoint reconstruction into ReFL and DRTune improves all five metrics on FLUX.1-dev with active sets and optimizer unchanged. Three instantiations evaluate the sampled image, eliminating connector residuals through Euler reconstruction (FlowBP-Sparse, FlowBP-Bridge) or targeting smaller residuals through multi-support quadrature (FlowBP-Lagrange). They retain activations for only a few velocity calls and at most one nested state Jacobian. Across SD3.5-M, FLUX.1-dev, and FLUX.2-Klein-base, they improve on prior direct-gradient baselines in preference metrics and overall compositional prompt following. Balancing alignment gains, training time, and tuning reliability, we recommend FlowBP-Sparse with K = 3 active calls as the default recipe.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.