Red-Teaming Chain-of-Thought Monitors by Compressing Hidden Scratchpads
Abstract
Chain-of-thought monitoring—having a trusted model review an AI actor's reasoning for signs of misbehavior—is one of the most promising defenses against misaligned AI systems. Evaluating the robustness of monitoring systems requires eliciting AI models to perform covert, misaligned actions that can be used to stress-test monitors. Thus far, existing methods (e.g. prompt optimization, finetuning) largely fail at eliciting successful evasion at a level greater than plain prompting. In this paper, we study the landscape of unmonitorability elicitations. We propose Scratchpad Compression, an RL-based method that achieves substantially higher successful evasion scores than all other elicitations we test, allowing for more effective red-teaming of misbehavior monitors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.