acceptodds
Under review as a conference paper at ICLR 2027

Before Reasoning Begins: Temporal Contrast Reveals a Causal Target Across Reasoning Tasks

Abstract

Code training is one of the most effective empirical routes to improving language-model reasoning, yet existing evidence remains largely at the level of data-mixture correlations and continued-training gains. A central mechanistic challenge is to identify an internal target that can be localized and causally tested across reasoning tasks. Our key insight is that reasoning-related intervention targets can be discovered before reasoning is generated. For the same task content, chat template and raw continuation inputs occupy different final pre-generation states even though both may later produce reasoning traces. We use this functionally aligned contrast to screen sparse autoencoder features at the final prompt position, control for responses reproduced by isolated prompt components, and localize the earliest consistent trajectory separation to layers L23–L24. This pipeline yields a compact set of 23 intervention targets in Qwen3-8B. Applying targeted SAE-guided suppression to these features throughout generation across 2,187 examples from GSM8K, CruxEval-O, and HumanEvalPlus reduces sample-weighted accuracy from \(86.1%\) to \(41.2%\), a drop of \(44.9\) percentage points. On GSM8K, same-size, layer-count-matched random null interventions reduce accuracy by only \(4.4\) points on average, compared with \(44.2\) points under the targeted intervention. Across the code benchmarks, targeted and random interventions exhibit task-dependent accuracy effects but distinct failure profiles. Visible chain-of-thought controls further show that removing visible reasoning does not reproduce the collapse, while forcing reasoning entry after intervention does not restore performance. These results identify a compact, high-leverage intervention target with broad causal influence across mathematical and code-related behavior. Together, these findings establish pre-generation contrast as a discovery strategy for identifying and testing causal targets across reasoning tasks. Our code is openly available at https://anonymous.4open.science/r/anonymous-reasoning-core-CEF4/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.