CARVE: Coordinating Imitation and Exploration through Reward Variance for LLM Reasoning Post-Training
Abstract
Reinforcement learning with verifiable rewards (RLVR) commonly follows supervised fine-tuning (SFT) in the post-training of reasoning models. However, the two stages often reuse the same problems even though group-relative RL depends on reward variation among rollouts. SFT can raise the solve rates of these problems and leave RL with many all-correct groups that provide no effective gradient, a failure mode we term variance starvation. Binary outcome rewards also cannot distinguish among correct derivations, while process supervision becomes less informative when trained on similarly saturated problems. To address these challenges, we propose CARVE, a variance-aware post-training framework with two coordinated designs. First, it separates SFT and RL problems and uses the fine-tuned parent to retain RL prompts with informative outcome variation. Second, it trains a process reward model on a subset of the screened problems, validates its selection ability before use, and restricts it to lowering the reward of weaker correct rollouts. The screened problems therefore provide both useful reward variation for policy optimization and informative supervision for process learning. Experiments across ten benchmarks and two backbone show that CARVE consistently outperforms shared-pool post-training baselines on mathematical reasoning while maintaining comparable performance on code and general reasoning. It achieves these gains with half the SFT data and substantially fewer RL updates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.