From Generating Reasoning to Relying on It: Direct Counterfactual Optimization for Chain-of-Thought
Abstract
A model can give a correct answer and a plausible chain-of-thought (CoT) even when that CoT has little causal influence on the answer. We examine whether training to improve answer accuracy or favor low-redundancy rationales reduces the CoT's causal redundancy with respect to the answer. We assess training effects through complementary behavioral and activation-based probes: CORD and all-layer activation patching, respectively. Outcome-based RLVR improves accuracy, and DPO changes generated CoTs, but neither consistently lowers held-out CORD. This leaves open whether causal redundancy is difficult to change or these training signals fail to reach the relevant answer behavior. We compare trajectory-level scalar rewards with direct gradients on counterfactual answer likelihood. Under the tested protocols, direct optimization substantially lowers CORD, whereas scalar credit does not consistently do so. Naive direct optimization, however, harms normal-task accuracy. We propose Direct Counterfactual Optimization (DCO), a training objective that discourages premature recovery of the original answer from insufficient reasoning while protecting the model's behavior on intact question-CoT paths. It uses a reference-relative target to limit suppression under early, question-masked CoT prefixes and a KL budget to allow limited changes on intact inputs. Across three held-out mathematical-reasoning sets, DCO achieves matched of to , with accuracy changes of to percentage points. All-layer patching shows that DCO and naive direct optimization reduce sensitivity to question replacement and increase sensitivity to CoT replacement, unlike DPO and auxiliary-reward RLVR.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.