Reasoning Gains and Capability Stability in Short-Horizon GRPO: A Paired KL Audit
Abstract
Does a reference-policy KL penalty protect capabilities during reasoning reinforcement learning, once target-task learning and training exposure are measured together? We compare six 300-update LoRA runs of Qwen2.5-1.5B-Instruct, pairing three seeds across KL coefficients zero and 0.02. Greedy accuracy on 300 held-out GSM8K questions rises from 58.7% to 68.7–72.3%. Meanwhile, likelihood accuracy remains at 77.7–78.2% on 600 BoolQ questions (77.8% initially) and 73.2–74.6% on 299 ARC-Challenge questions (74.6% initially). The paired KL effects on final BoolQ change are , , and percentage points. Thus this coefficient has no consistent protective effect across the three seeds. Item-level predictions distinguish aggregate stability from preservation of individual decisions: two KL runs finish at the initial BoolQ accuracy despite four or five gains canceling the same number of losses. An exact enumeration of compatible pairings shows how much statistical evidence is lost when only accuracy totals are retained. We also specify why a penalty measured on arithmetic responses does not bound changes on another prompt distribution. The evidence supports substantial arithmetic learning with small changes on two likelihood tasks in this short-horizon setting, while leaving broader preservation and stronger KL interventions unresolved.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.