acceptodds
Under review as a conference paper at ICLR 2027

On Reasoning Correctness Directions

Abstract

We present experimental evidence suggesting that information about whether a model’s reasoning is on the right track can be linearly decodable without defining a causally useful direction for steering. To study this, we introduce ChainSearch, a controlled arithmetic reasoning task that extends Countdown with explicit search and backtracking. We train a 3-million-parameter model to solve this task. Its compact size allows us to study internal model correctness directions in detail using linear probes, mean activation differences, and direct optimisation, as part of a rigorous search for internal reasoning-correctness directions. We find that while the model contains statistically significant internal directions of reasoning correctness, none of the resulting directions consistently improves reasoning under causal intervention. This calls into question whether the linear representation hypothesis is sufficient for identifying causally useful directions in the context of reasoning. Our code is made available together with this work.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.