FlowSAPO: Semantics-Aware Regional Optimization for Visual Generation
Abstract
Online reinforcement learning aligns diffusion and flow models using rewards evaluated on complete images or videos, while policy optimization acts on denoising transitions over a latent grid. A single global clipping statistic discards the spatial distribution of local likelihood changes, whereas clipping each latent position separately does not coordinate related positions. We introduce FLowSAPO, which combines exact positionwise likelihood ratios with regional surrogate weights constructed from the model's own text attention. Connected, overlapping regions are built during rollout and frozen during replay. Within each region, local log ratios are aggregated with inverse coverage weights and exponentiated before clipping, while the reward and scalar advantage remain unchanged. Under Gaussian transitions with shared covariance at a fixed state, we show that the global statistic depends on local divergences only through their total. We also show that the regional surrogate gradient at the rollout policy is proportional to the local gradient for any stored cover; away from that point, the cover affects both exponential weights and shared clipping decisions. Across the reported image and video comparisons, FLowSAPO improves over the local objective, including GenEval attribute binding from 0.86 to 0.96 and All Frame OCR on Wan2.1 from 0.339 to 0.427. It also raises VideoReward visual quality from 2.302 to 3.901 and motion quality from 0.811 to 1.596 compared with flowgrpo. These results support regional optimization as a useful complement to positionwise clipping for visual generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.