Greedy-Prompt Policy Gradient Flow
Abstract
Reinforcement learning with verifiable rewards (RLVR) optimizes a policy over a prompt distribution. Inspired by recent findings that RLVR remains effective when only one or a few prompts are used for gradient updates, we study how each prompt gradient shapes policy-optimization dynamics. This perspective leads to a greedy flow that selects the prompt maximizing instantaneous improvement. Although effective, one can hypothesize that the advantage has a trivial explanation: the selected update may simply have a larger component along the population gradient. We therefore remove this parallel advantage through normalization and find that the normalized flow remains effective, showing that parallel rescaling alone does not explain the observed gap in LLM RLVR. Our diagnostics further document substantial orthogonal motion and nontrivial temporal geometry. Motivated by these observations, we study projection-steering regularity as one possible sufficient condition for global dominance over standard gradient flow. Empirically, we design algorithms approximating these flows and evaluate them on Qwen3-8B-Base with LoRA over DAPO-Math-17K. The results show consistent dominance throughout the observed training horizon, with gains transferring across multiple out-of-distribution benchmarks, including MATH-500, AMC 2023, AIME 2024, AIME 2025, Minerva Math, and OlympiadBench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.