acceptodds
Under review as a conference paper at ICLR 2027

Behavior Cue Reasoning: Monitorable Reasoning Improves Safety and Efficiency through Oversight

Abstract

Reasoning in Large Language Models (LLMs) poses a challenge for oversight as many misaligned behaviors can surface only in downstream actions once reasoning has concluded. To address this, we introduce Behavior Cue Reasoning for making LLM reasoning more controllable and monitorable. Behavior Cues are special token sequences that a model is trained to emit immediately before specific implicit and explicit behaviors, acting as dual purpose signal and control levers. Our experiments reveal that a Behavior Cue Reasoning model has equal to greater performance as the base model, allows for steerable reasoning through external enforcement of Behavior Cues, and improves the monitorability of reasoning for external oversight monitors. When leveraged by an oracle monitor in an environment where excessive constraint violations results in failure, Behavior Cues allows for the recovery of safe actions from up to 83% of reasoning traces that would otherwise end with the proposal of an unsafe action. When applying early-stopping efficiency methods, Behavior Cues see a similar token-saving, accuracy trade off without the need for otherwise costly answer probes. Through evaluation across two model families and three domains, we show that Behavior Cue Reasoning improves reasoning monitorability and controllability with no cost to performance. More broadly, our work progresses scalable oversight by demonstrating how the monitored model itself can be trained to reason in a more tractable to oversight.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.