acceptodds
Under review as a conference paper at ICLR 2027

Competing Rewards for Exploration in Safety

Abstract

Safety training on language models often induces over-refusal: improved safety on harmful prompts at the cost of increased refusal on harmless ones. We argue that this problem arises because of limited exploration within safety reasoning. Current models often align their reasoning with the final safety objective, avoiding unsafe reasoning altogether or failing to produce a safe answer after unsafe interpretations. In this paper, we address the safety-refusal trade-off with our key insight is that unsafe reasoning can itself serve as a useful exploratory signal. Rather than preemptively blocking harmful thoughts, we encourage the model to sufficiently explore unsafe reasoning but produce a safe response. The harmful exploration improves the model's ability to distinguish harmful from harmless prompts by resolving ambiguity, allowing it to remain safe while complying only when appropriate. We cast this as an adversarial optimization problem in which a reasoning player explores strategies for producing an unsafe response and an answer player ensures a safe final output. We train a single model with dense rewards to play both roles within one chain-of-thought, across different segments. To achieve this, we find that process rewards are crucial for stable optimization of competing objectives. Our resulting model, SEAR, deliberately engages in harmful reasoning as exploration, improves the safety-refusal trade-off on over-refusal benchmarks, and is more robust against attacks that directly manipulate the reasoning to be harmful. Our results show that competing rewards provide an effective way to encourage exploration in safety reasoning without requiring the reasoning process itself to follow the final safety objective.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.