Rewarding the Wrong Circuit: How RLVR Reshapes Safety in Reasoning Models
Abstract
Reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning, yet most safety evaluations compare only endpoints measured before and after training. We ask a mechanistic question that endpoint scores cannot answer: does safety remain in control after reward optimization? We track the same aligned model before, during, and after RLVR focused on capability and separate two functions that are usually conflated: recognizing that a request is unsafe and allowing that recognition to govern the final policy. Our hypothesis is decoupling between recognition and policy: RLVR can preserve harm recognition while reallocating causal control toward answer commitment driven by reward. We identify Reward Override Components (ROCs) by a dual causal criterion: a component must gain mediation of verifiable reward while losing mediation of refusal. On the frozen trajectory from Qwen2.5-7B-Instruct to Code-R1, the gap between recognition and refusal increases from 19.3 to 36.6 points; restoring ROC states from the reference checkpoint recovers 24.6 pp refusal. We then introduce Circuit Anchor RLVR (CA RLVR), which anchors only this discovered interface during capability RLVR. Under the protocol that follows the same model lineage, CA RLVR achieves 17.4% harmful compliance on reasoning jailbreaks versus 29.8% for the strongest adapted baseline while retaining 62.2 code average versus 62.5 for unmodified RLVR.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.