acceptodds
Under review as a conference paper at ICLR 2027

MEND: RL VIA PROXIMAL TARGET MATCHING

Abstract

Reward-based post-training of flow models often relies on KL penalties, refer- ence models, or thousands of updates. Direct reward backpropagation provides a direction but does not check whether each proposed change justifies its magni- tude. We introduce MEND, a reinforcement learning method based on proximal target matching. A reward cap assigns unchanged targets to high-scoring samples. Below the cap, MEND proposes moves along the reward gradient and accepts a move only when its capped reward gain exceeds a quadratic displacement penalty. The model learns the accepted displacements through regression, without a KL term, frozen reference model, or advantage weights. MEND scores higher than the Flow-GRPO adapter on five of six evaluators at the same reported distance to base-model images. Under an equal-budget protocol, it exceeds the reported ReFL and DiffusionNFT curves at every reported update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A three-reward run also exceeds the released DiffusionNFT checkpoint on all three training re- wards. The method transfers to another flow backbone and a distilled model. Code and checkpoints will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.