acceptodds
Under review as a conference paper at ICLR 2027

Pressure Is Not Enough: A Causal Test of Whether Optimization Induces Encoded Chain-of-Thought

Abstract

If a chain-of-thought (CoT) monitor penalizes certain words, does optimizing against that penalty cause a model to develop an alternative, causally load-bearing representation that the monitor does not recognize? We test this on Coin Flip, a synthetic state-tracking task, using GRPO. We first separate two confounds. Across several discovery and adoption mechanisms, we find no reliable spontaneous non-literal encoding. Separately, an SFT-seeded Llama-3-8B-Instruct checkpoint demonstrates that the capability exists: bidirectional counterfactual interventions succeed in 42/42 trials before RL and 167/168 trials across four seeds after 150 GRPO steps. We then resume RL from a genuinely literal checkpoint and apply lexical pressure from the first step. The penalty fires on every sampled main rollout throughout 150 steps, yet no alternative representation emerges: final accuracy shows no statistically detectable difference from an unpenalized control (p=1.0), and all 160 final-checkpoint completions and 9,600 training rollouts remain literal. Thus, under the tested conditions, lexical pressure did not induce an alternative representation despite verified capability to use one when supplied. We additionally identify two apparent discoveries that fail stricter causal tests and find an initial semantic probe of monitor evasion to be inconclusive.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.