Holding the Per-Step KL: Why Clipping Cannot, but Control Can
Abstract
GRPO-based reinforcement learning (RL) aligns flow-matching models with task rewards, but ratio clipping does not directly control the KL divergence of policy updates. We analyze this gap through an exact Gaussian decomposition of the per-step log-ratio, in which the per-step KL sets both its mean and variance. This reveals a distinction between calibrating the clipping frequency and controlling policy displacement: for the centered threshold rules considered, standardization removes the first-order dependence of the clipping gate on the overall displacement scale. Building on this distinction, we develop the method in three layers. F-Clip clips the centered log-ratio in KL-scaled coordinates; Z-Clip adapts the clipping range to the batch spread, so that its threshold corresponds to a target clipping fraction; and K-Hold uses learning-rate feedback to keep the floor-corrected mean KL over selected denoising steps near a prescribed target. Our analysis provides a unified characterization of ratio clipping in flow-matching RL, connecting the per-step log-ratio to the KL quantity that must ultimately be controlled. Experiments on OCR-based text rendering with SD3.5-M and FLUX.2-klein validate the predicted clipping and KL behavior. K-Hold improves held-out reward-model scores over the baselines, while preserving the gain under matched OCR performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.