Chain-of-Thought Reliance does not Guarantee Monitorability
Abstract
A prevailing hypothesis in chain-of-thought (CoT) monitorability, the ability of a monitor to determine whether a model is taking an undesirable action, is that models are more monitorable when working on difficult tasks. Simply increasing task difficulty increases a model's reliance on its CoT, which in turn makes the CoT more monitorable because it can't be used for post-hoc rationalization. We find that this optimistic view often does not play out in practice. This paper presents the first systematic, empirical study of how monitorability varies with task difficulty and CoT reliance. In hint-based monitorability evaluations, which monitor for a model's usage of a biasing input feature, increasing CoT reliance does not improve monitorability, and can even reduce it by up to 30 percentage points. In some settings, this decrease is correlated with models generating longer CoTs for more difficult tasks; we propose a simple truncation heuristic that recovers up to 16 percentage points. Increased CoT reliance proves more helpful in secret-side-task evaluations, which adversarially prompt the model to complete a side task without the monitor noticing. There, increased CoT reliance eliminates a model's ability to complete the side task without verbalization. We observe other strategies for monitor evasion that are not diminished by increased CoT reliance. Overall, while increased CoT reliance can improve the narrow class of monitor errors caused by non-verbalization, it alone does not guarantee CoT monitorability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.