Learning from Sparse Successes via Balanced Group Reconstruction
Abstract
Group-relative reinforcement learning depends on finding informative combina- tions of successful and unsuccessful responses. When successes are scarce, a verified sibling can seed further sampling. We introduce Success-Conditioned Reconstruction and Augmentation (SCRA): generate additional solutions from a short successful-sibling prefix, verify them, and reconstruct the maximum-available 1:1 group for training on the original no-hint question, then finish with ordinary GRPO. On Qwen3-8B, SCRA reaches 90.04 / 39.17 / 31.25 on MATH-500, AIME 2024, and AIME 2025 over three training seeds, and 71.48 on chemistry. It is strongest on both AIME suites and on every SciKnowEval split; HiPO remains slightly ahead on MATH-500. SCRA stays above GRPO on every reported split, including a GRPO run with 16 original samples that matches the extra sampling budget of triggered SCRA steps on the mathematical protocol. A same-protocol SDPO run collapsed on the mathematical suites after entropy dropped, but stayed above the base model on the four SciKnowEval splits. On the same full-benchmark protocol, 1:1 reconstruction is strongest on both AIME suites among the tested allocations, extra no-hint sampling does not match the sibling prefix on AIME, and withdrawing the scaffold after 81 of 135 steps outperforms a later switch.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.