acceptodds
Under review as a conference paper at ICLR 2027

Red-Teaming Chain-of-Thought Monitors by Compressing Hidden Scratchpads

Abstract

Chain-of-thought monitoring—having a trusted model review an AI actor's reasoning for signs of misbehavior—is one of the most promising defenses against misaligned AI systems. Evaluating the robustness of monitoring systems requires eliciting AI models to perform covert, misaligned actions that can be used to stress-test monitors. Thus far, existing methods (e.g. prompt optimization, finetuning) largely fail at eliciting successful evasion at a level greater than plain prompting. In this paper, we study the landscape of unmonitorability elicitations. We propose Scratchpad Compression, an RL-based method that achieves substantially higher successful evasion scores than all other elicitations we test, allowing for more effective red-teaming of misbehavior monitors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.