The First Few Tokens Matter: Understanding Chain-of-Thought Controllability
Abstract
Chain-of-thought (CoT) monitoring is a promising way to oversee increasingly capable agents, in part because current models appear limited in their ability to control what they verbalize in their reasoning traces. We investigate the source of this constraint and link it to an entropy collapse in the first few tokens following the opening thinking tag. We exploit this mechanism to recover CoT controllability through two inference-time interventions on initial tokens: prefilling, where even a single token can substantially increase controllability, and banning frequent opening tokens combined with in-context examples, which improves average compliance by 34-41 percentage points across models. These interventions recover mainly addition and style controllability (adding or transforming information), while suppression controllability (avoiding or removing information) exhibits more varied responses, suggesting distinct underlying mechanisms. Consistent with this, the two groups behave as separate capabilities: across models, performance correlates positively within these groups, while cross-group correlations are mostly negative. Importantly, we find that associations with monitorability depend on the evaluation: addition and style controllability more consistently correlate positively with monitorability, while suppression sometimes shows negative associations. Aggregate controllability can therefore obscure these relationships, motivating separate reporting of control tasks and categories when evaluating CoT monitorability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.