acceptodds
Under review as a conference paper at ICLR 2027

Greedy-Prompt Policy Gradient Flow

Abstract

Reinforcement learning with verifiable rewards (RLVR) optimizes a policy over a prompt distribution. Inspired by recent findings that RLVR remains effective when only one or a few prompts are used for gradient updates, we study how each prompt gradient shapes policy-optimization dynamics. This perspective leads to a greedy flow that selects the prompt maximizing instantaneous improvement. Although effective, one can hypothesize that the advantage has a trivial explanation: the selected update may simply have a larger component along the population gradient. We therefore remove this parallel advantage through normalization and find that the normalized flow remains effective, showing that parallel rescaling alone does not explain the observed gap in LLM RLVR. Our diagnostics further document substantial orthogonal motion and nontrivial temporal geometry. Motivated by these observations, we study projection-steering regularity as one possible sufficient condition for global dominance over standard gradient flow. Empirically, we design algorithms approximating these flows and evaluate them on Qwen3-8B-Base with LoRA over DAPO-Math-17K. The results show consistent dominance throughout the observed training horizon, with gains transferring across multiple out-of-distribution benchmarks, including MATH-500, AMC 2023, AIME 2024, AIME 2025, Minerva Math, and OlympiadBench.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.