Precision Bounds Diversity: Reward Engineering for Instruction Following via RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving instruction following (IF), with recent work assuming that scaling constraint diversity is essential for generalization. We challenge this assumption and show that it is fundamentally misguided. By decoupling training data into hard-only, soft-only, and mixed subsets, we find that models trained exclusively on code-verifiable hard constraints outperform those trained on LLM-judged soft constraints on average, including on soft-constraint benchmarks, and match mixed training despite using fewer constraint types. Diagnosis reveals that LLM-based reward models suffer from critically low recall and a systematic inability to penalize violations, with reliability degrading sharply as constraint count increases. Controlled noise-injection experiments further confirm that reward precision, not constraint diversity, is the primary driver of RLVR performance. Motivated by these insights, we propose H1S, a data-composition strategy that retains all hard constraints while pruning soft constraints to at most one per instruction, confining LLM evaluation to its most reliable single-constraint regime. On five IF benchmarks, H1S achieves average gains of 11.3% (7B) and 6.7% (32B) over base models, surpassing mixed-constraint training by 5.5% at 7B while reducing per-step training latency by 58% and preserving general capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.