Adaptive Contrastive Self-Play for Regime-Aware Reinforcement Learning Fine-Tuning of Language Models
Abstract
Reinforcement learning with verifiable rewards (RLVR) relies on variation among sampled rewards. When every response in a rollout group receives effectively the same shaped verifier reward, Group Relative Policy Optimization (GRPO) receives no policy gradient from reward for that group. We ask whether the optimization objective should change as such groups become common. We introduce Adaptive Contrastive Self-Play (A-CSP), which monitors the fraction of zero variance groups and switches between GRPO and a contrastive objective that learns from higher and lower reward responses when eligible pairs exist. Across 156 training runs on GSM8K and a custom MATH-L1-3 pool, we find that the value of switching is strongly regime dependent. On Qwen2.5-1.5B GSM8K, A-CSP reaches similar peak exact match to GRPO, 0.730 versus 0.737, while retaining more of its peak performance at the end of training, 0.97 versus 0.90. On Qwen2.5-3B, A-CSP improves peak accuracy on both tasks and final accuracy on GSM8K, while final accuracy on MATH-L1-3 is essentially unchanged. Mechanistic controls show that masking KL on zero variance groups does not reproduce the A-CSP behavior, while the decomposition suggests that batch composition affects optimization dynamics more than KL masking alone on the flagship setting. Simple switching schedules can also match or exceed A-CSP in peak accuracy, showing that zero variance alone is not sufficient to determine the best switching policy. These results show that adaptive objective switching can improve RLVR training in some regimes, but its benefit depends on more than the prevalence of zero variance groups.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.