acceptodds
Under review as a conference paper at ICLR 2027

Safety and Reasoning Vulnerabilities in Self-Rewarding Test-Time Reinforcement Learning

Abstract

Test-time reinforcement learning adapts a language model on unlabeled test inputs by deriving rewards from the model's own generations. This label-free property makes test-time RL attractive when verifiable feedback is unavailable, and recent methods have been shown to improve reasoning when the test-time stream contains only reasoning problems. We study what happens in heterogeneous streams, i.e., when reasoning problems are mixed with jailbreak attempts or benign instructions. We evaluate two representative methods: TTRL, which rewards majority-vote agreement, and RENT, which rewards low entropy. We find that mixed streams produce 3 linked effects: safety amplification, harmfulness amplification, and a reasoning tax. These effects follow a common mechanism: self-rewarding test-time RL reinforces the response pattern that its reward can score most easily. Under TTRL, jailbreak prompts often collapse to a short refusal template; on held-out jailbreak prompts, 54-78% of responses can become the same refusal sentence. This lowers the measured attack success rate, producing apparent safety amplification, while math responses become more self-consistent without becoming more correct, producing the reasoning tax. RENT does not show the same refusal-template collapse, and its safety direction depends more strongly on the model's initial behavior. We further show harmfulness amplification with an adversarial prompt, "HarmInject”, that couples a jailbreak to a reasoning problem, enabling rewards for harmful compliance. These results show that consistency and confidence are unsafe proxies for correctness and safety in heterogeneous test-time RL.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.