Improving Chain-of-Thought Monitorability in Large Language Models: Causal Bounds and Post-Training
Abstract
Chain-of-thought (CoT) monitoring aims to detect potentially harmful behaviors in language models from their reasoning traces. Intervention-based evaluations measure this property by asking whether a monitor can detect the role of an intervention from the CoT when the targeted behavior occurs, commonly using minimal-criterion sensitivity () guan2025monitoring. However, a target outcome observed under intervention may also have occurred without it. For stable binary intervention experiments with a positive average effect, we show that upper-bounds the monitor's recall on intervention-caused outcomes and derive a corresponding lower bound. We then study whether post-training improves measured CoT monitorability under a fixed monitoring setup and whether the observed gains persist when sensitivity is replaced by this lower bound. We construct a training dataset of paired control–intervention tasks spanning contextual evidence, answer cues, sandbagging, and deception, and use it to train models with supervised fine-tuning, reinforcement learning with monitoring-based rewards, and on-policy self-distillation. Across most evaluated tasks, all three methods improve measured monitorability relative to the base model under both the standard composite and its causal-recall lower-bound counterpart.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.