Same Mastery, Opposite Directions: Signed Learning Progress for Prompt Revisiting in RLVR
Abstract
While reinforcement learning with verifiable rewards has improved reasoning, the dominant cost is grouped on-policy generation, and most methods still spend a fixed group budget according to a prompt's current state. We study this allocation through learning progress, the recent change in pass rate relative to a slower mastery state. A comparison of uniformly sampled GRPO revisits uncovers two facts. First, current difficulty, competence match, and the magnitude of historical advantage assign the same priority to two prompts with the same pass rate, whether the policy is improving or regressing, while an all-correct or all-incorrect group has zero relative advantage. Second, with mastery and the size of the stored change held fixed, the next visit raises the pass rate by 22.4 points when that change is positive and by 1.1 points when it is negative, so another group is better spent on the improving prompt. Based on these observations, we propose LP-GRPO, which reads a slow mastery state and a fast signed progress state from before the current group. The state is used twice: LP-Sampling spends the next groups where a nonzero advantage is still available and progress is positive, and LP-Reweighting multiplies the resulting nonzero advantages by a positive weight from the same state. Neither step adds a selector or a value model. On eight multimodal benchmarks, with the same number of retained groups per update, LP-GRPO raises GRPO from 66.50 to 69.75 and exceeds the strongest matched baseline by 1.62 points. The revisit rule alone reaches 69.18. The full method matches GRPO's final average in half the updates, and CDAS and PSPO in two thirds.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.