acceptodds
Under review as a conference paper at ICLR 2027

Chain of Thought Monitoring Can Fail Even When CoT Is Necessary

Abstract

Chain of Thought (CoT) monitoring is used by frontier labs to detect misbehaviour of AI systems. Yet, its reliability remains poorly understood. We attribute this to the lack of a successful organizing principle that leads to generalizable conclusions about monitoring reliability in different regimes. Recent work attributes monitoring failures to the CoT not being necessary for the answer, and argues that CoT becomes monitorable when it is necessary. We show that this binary necessary/not-necessary organizing principle is too coarse by demonstrating that monitoring failures persist even if the CoT is necessary for the answer. In response, we introduce a more fine-grained organizing principle based on a hierarchy of four CoT necessity levels. Building on this hierarchy, we develop a framework for measuring CoT monitorability at the three lowest levels via intervention experiments. We demonstrate monitoring failures at all three of these levels on realistic, hard prediction tasks. On a difficult medical-diagnosis task, patient sex and a clinical attribute influence the answer in at least 19% and 25% of correct responses with a necessary CoT, yet neither influence is ever articulated. These results show that even when CoT is necessary, it is not guaranteed to be monitorable.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.