Practice, Correct, and Commit: Uncertainty-Aware Checkpoint Commitment for Tutor-Guided RLVR
Abstract
Adopting a teacher-corrected checkpoint changes the model that generates subsequent training data in reinforcement learning with verifiable rewards (RLVR), making checkpoint adoption a consequential training decision. We separate producing a corrected candidate from deciding whether to use it for further learning. UQ-Commit implements this separation by retaining the checkpoint after each RL block, constructing a teacher-corrected candidate, and comparing the two before training continues. It adopts correction when the lower bootstrap quantile of the paired validation pass@32 difference exceeds a fixed margin. In an eight-block study with a 4B student and three training seeds, UQ-Commit reaches 53.06 Macro6, the mean pass@32 percentage across six mathematical-reasoning benchmarks. The reported mean gain is 2.25 percentage points over always-accept correction under the same teacher, correction recipe, and GRPO schedule, and 2.18 points over GRPO alone. Separate four-block 4B and 8B comparisons provide additional positive mean gains over GRPO. These results establish selective checkpoint adoption as a useful component of the evaluated teacher-correction pipeline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.