DO REASONING MODELS EVEN START? DECOMPOSING COT CONTROLLABILITY
Abstract
The headline metric for chain-of-thought (CoT) controllability—the fraction of reasoning traces that comply *perfectly* with an instruction—conflates three quantities and is blind to one of them: how well a model *maintains* a constraint once it has started. Low CoT controllability is a pillar of the safety case for CoT monitoring, so *why* it is low decides whether that case is robust or one design choice away from failing. We grade 272k reasoning traces with position-resolved graders and treat the first violation as a time-to-event outcome, separating onset (does the model ever adopt the constraint), maintenance (slips per 1,000 words once adopted) and exposure (trace length), and we intervene on onset by opening the reasoning with a short compliant prefix. Across six open-weight models from three families (Qwen3, OLMo-3-Think, gpt-oss), the uppercase and alternating-case constraints are never started in five: every uppercase trace fails at the first word, and constrained traces carry the style no more often than unconstrained ones; gpt-oss-120b starts uppercase in half of its traces and drops it within a few words, and the two gpt-oss models start the required opening string but rarely close with it. Keyword suppression, in contrast, has no onset failure and leaks from the first word at a rate set by how often the question calls for the word. The benchmark score for "reason in uppercase" is 0% for every model we test, yet after a 20-word compliant prefix the slip rates of Qwen3-8B/32B and OLMo-3-7B/32B span 13× (95% CI 10–17), improve with scale within two of three families, and Qwen3-32B keeps an entire trace in uppercase 38–45% of the time, largely by continuing its context rather than by following the instruction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.