Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training
Abstract
Self-improvement can self-regress. In verifiable-reward code RL, continued optimization can first improve a language model and then erase those gains. We study this instability with Qwen-2.5-3B/7B on custom competitive-programming splits drawn from HumanEval, MBPP, and APPS, using binary CodeGrader reward and chained training campaigns. In a 200-step Qwen-2.5-7B REINFORCE diagnostic run, the online success signal rises from 25% to 81% within roughly 50 updates and later falls to nearly zero; checkpoint-continuation tests show that checkpoints from different parts of this trajectory lead to sharply different downstream outcomes. We therefore ask an operational question: When does regression become actionable, and where should the control loop intervene? We separate two temporal quantities: a within-campaign cliff, measured by the loss from a campaign's online peak to its endpoint, and cross-campaign endpoint persistence, measured from the sequence of checkpoints passed forward through the chain. The distinction exposes a striking asymmetry. On the matched 7B held-out splits, the frozen base reaches 10.4% [8.1, 12.8], vanilla REINFORCE 11.8% [5.2, 18.3], and GRPO 20.7% [15.7, 25.1] over five seeds, yet the mean within-campaign peak-to-end gap remains nearly unchanged under REINFORCE and GRPO (17.6 vs. 16.5 points). A matched REINFORCE+leave-one-out-baseline (RLOO-style) control reaches 13.0% [8.5, 17.5] while retaining a 17.0-point mean peak-to-end gap. GRPO raises the floor, but does not remove the cliff. The RLOO control further shows that this endpoint gain is not reproduced by adding a simple baseline alone. Control effectiveness is also regime-dependent. A simple campaign-end controller raises 3B end-of-chain pass@1 from 4.9% to 9.5% across five matched seeds (paired bootstrap difference [+0.4,+8.9] points), but provides no clear gain at 7B. A trajectory-informed peak-rollback/budget schedule reaches 22.2% [14.1,28.0] at 7B (). Turning the causal-prefix three-decline rule into a live controller raises 7B held-out end-of-chain pass@1 to 25.6% [20.3,30.4] over five seeds and reduces the mean peak-to-end gap from 17.6 to 6.2 points. The central lesson is that final performance is not a sufficient statistic for training stability: update-rule changes can raise the floor while same-campaign control directly suppresses the cliff. We advocate reporting local peak retention, failure geometry, endpoint persistence, and detection lead time alongside final pass@1 when evaluating self-improving RL systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.