Chain of Thought Monitoring Can Fail Even When CoT Is Necessary
Abstract
Chain of Thought (CoT) monitoring is used by frontier labs to detect misbehaviour of AI systems. Yet, its reliability remains poorly understood. We attribute this to the lack of a successful organizing principle that leads to generalizable conclusions about monitoring reliability in different regimes. Recent work attributes monitoring failures to the CoT not being necessary for the answer, and argues that CoT becomes monitorable when it is necessary. We show that this binary necessary/not-necessary organizing principle is too coarse by demonstrating that monitoring failures persist even if the CoT is necessary for the answer. In response, we introduce a more fine-grained organizing principle based on a hierarchy of four CoT necessity levels. Building on this hierarchy, we develop a framework for measuring CoT monitorability at the three lowest levels via intervention experiments. We demonstrate monitoring failures at all three of these levels on realistic, hard prediction tasks. On a difficult medical-diagnosis task, patient sex and a clinical attribute influence the answer in at least 19% and 25% of correct responses with a necessary CoT, yet neither influence is ever articulated. These results show that even when CoT is necessary, it is not guaranteed to be monitorable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.