acceptodds
Under review as a conference paper at ICLR 2027

MoRe: Sensing Prompt Utility Through Gate Modulation for Sample-Efficient RLVR

Abstract

Reinforcement learning with verifiable rewards (RLVR) must allocate a limited rollout budget across prompts. We observe systematic differences in full-parameter gradient energy across prompts, even after accounting for observed success counts and prompt length. To characterize one model-internal source of this variation, we introduce modulation reactivity (MoRe), a closed-form measure of final-layer prompt-cache sensitivity to gate modulation that preserves bilinear interactions. MoRe uses only current weights and prefill activations, requiring no additional rollouts or backward passes. The same Jacobian connects cache responses to gradient feedback, giving MoRe a normalized expected gradient-energy interpretation under isotropic reference feedback. Among prompts matched on observed success rate and length, the high-MoRe group exhibits up to twice the cache-path feedback energy of the low-MoRe group. Fixing within-prompt attention proportions removes more than half of this matched feedback gap without changing MoRe. These findings motivate a two-stage sampler that uses reward beliefs to retain prompts likely to yield mixed rewards, then allocates a fixed number of rollout slots using current MoRe scores. On mathematical reasoning benchmarks, the sampler improves average accuracy by 1.92 percentage points over vanilla GRPO.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.