Robustness Is Not Reasoning: A Double Dissociation in Latent Chain-of-Thought
Abstract
Latent chain-of-thought methods such as Coconut and CODI replace token-level reasoning with continuous "thoughts," and a growing literature probes whether these thoughts are robust to perturbation and whether they carry the reasoning, often collapsing both questions into a single perturbation test. We show that these are two distinct properties that can be dissociated. We apply three interventions to individual latent thoughts: random perturbation, counterfactual patching, and content-shuffle (replacing a thought with a real thought computed for a different input). We find that a thought's robustness to random noise is largely independent of whether its content is necessary for the answer. On a controlled k-hop modular-arithmetic task (Pythia-160m/410m, multi-seed, accuracy-matched), the fraction of thoughts that are noise-inert rises monotonically with recurrence depth (from 0.00 at depth 3 to 0.87 at depth 9), while the content of those same thoughts remains necessary at moderate depth and only degrades at excess depth, so the "robust and load-bearing" regime is non-monotonic. On a released CODI-Llama-1B model on GSM8K, latent thoughts are simultaneously noise-robust (answers rarely flip under perturbations of 16–64× the thought norm), individually redundant (single-thought ablation costs at most 4.3% accuracy), collectively essential (ablating all thoughts returns accuracy to near the no-thought floor), and content-load-bearing (shuffling content collapses accuracy from 0.436 to 0.165 ± 0.004), together operating as a distributed, error-correcting code that is individually redundant yet collectively essential. Robustness is therefore not evidence of reasoning, and robustness-only diagnostics can misrepresent what latent thoughts do.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.