Which Updates Erase What the Base Model Knew? Optimizer-Conditioned Attribution of Reasoning-Support Loss in RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) can improve short-horizon accuracy while erasing solution routes already supported by the base model. Entropy decline reveals this contraction but does not identify which updates cause it. We introduce CAGE (Cross-context Attribution of Gradient Effects), an optimizer-conditioned cross-context attribution score that estimates how each current update changes entropy at held-out correct states and selectively regularizes the predicted high-risk tail. Accumulated CAGE exposure predicts subsequent support loss () and remains predictive when recent entropy decline is included on the same rows. CAGE improves macro Pass@1 by +2.06 pp over OPEFO. Under the construction-matched discrete intervention, CAGE-D exceeds matched random targeting by +2.59 pp in macro Pass@1 and +.037 in correct-family mass, and exceeds OPEFO by +0.93 pp in macro Pass@1; uniform KL, source-strength matching, and mean-probe controls recover only parts of the gain. Selective KL recovers most of the high- loss, with CAGE-D above matched random selection in point estimate. It costs GRPO wall-clock and yields positive transfer point estimates on Qwen-7B, Llama-8B, and code, with weaker accuracy evidence under the larger shifts. Together, these results support targeting update-level risk rather than treating entropy preservation alone as sufficient for intervention design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.