acceptodds
Under review as a conference paper at ICLR 2027

NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

Abstract

Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that are not captured by the reward proxy. We identify a simple structural signature of this drift: across three post-training methods (NFT, AWM, DPO), RL fine-tuning inflates the per-step velocity norm by to relative to the reference. This inflation is perceptually consequential: because the generated sample is the displacement accumulated by the velocity field, a norm inflation raises the second moment of the sample by roughly , which surfaces in the decoded image as deviations in luminance and chrominance, i.e., boosted contrast, over-sharpened edges, and unnatural white balance. The obvious remedy, rescaling back to at inference time as done for classifier-free guidance, does not transfer to RL: it neither improves reward nor fixes the quality degradation, because the inflation is co-adapted into the model weights. We further show that velocity magnitude rescaling carries no coherent reward signal at the batch level, indicating that suppressing norm inflation is unlikely to remove a consistently reward-carrying component. Both findings point to NormGuard, a simple hinge penalty that activates only when exceeds and composes additively with any velocity-local base loss. Across two base models, three post-training methods, and two reward proxies, NormGuard consistently improves MLLM-judged image quality, usually improves forensic realism, and largely preserves reward, with gains that amplify under few-step inference and are not explained by early stopping.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.