acceptodds
Under review as a conference paper at ICLR 2027

Rewarding the Wrong Circuit: How RLVR Reshapes Safety in Reasoning Models

Abstract

Reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning, yet most safety evaluations compare only endpoints measured before and after training. We ask a mechanistic question that endpoint scores cannot answer: does safety remain in control after reward optimization? We track the same aligned model before, during, and after RLVR focused on capability and separate two functions that are usually conflated: recognizing that a request is unsafe and allowing that recognition to govern the final policy. Our hypothesis is decoupling between recognition and policy: RLVR can preserve harm recognition while reallocating causal control toward answer commitment driven by reward. We identify Reward Override Components (ROCs) by a dual causal criterion: a component must gain mediation of verifiable reward while losing mediation of refusal. On the frozen trajectory from Qwen2.5-7B-Instruct to Code-R1, the gap between recognition and refusal increases from 19.3 to 36.6 points; restoring ROC states from the reference checkpoint recovers 24.6 pp refusal. We then introduce Circuit Anchor RLVR (CA RLVR), which anchors only this discovered interface during capability RLVR. Under the protocol that follows the same model lineage, CA RLVR achieves 17.4% harmful compliance on reasoning jailbreaks versus 29.8% for the strongest adapted baseline while retaining 62.2 code average versus 62.5 for unmodified RLVR.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.