acceptodds
Under review as a conference paper at ICLR 2027

Internalizing the Halt Vector: Full-Vector and On-Axis Reconstruction

Abstract

Can a reasoning model learn earlier stopping by matching only the projection of a steered activation, rather than reconstructing the full vector? We compare full-vector and on-axis reconstruction on DeepSeek-R1-Distill-Qwen-7B, separating reasoning length, mathematical answer correctness and compliance with the requested answer format. An outcome-informed follow-up to an earlier held-out comparison, with development-only selection, evaluated three shortening targets on 2,000 held-out problems. One target established matched shortening within one percentage point: full-vector and on-axis reconstruction reduced thinking by 23.56% and 23.39%, respectively. Their accuracy was equivalent within the prespecified 2.6-point margin under Math-Verify, while a boxed-answer grader favored full-vector reconstruction by 2.80 points. The on-axis model produced more closed, unboxed answers despite an explicit boxing instruction. A post-hoc, model-assessed audit supports a predominantly formatting-based explanation of this discrepancy, subject to assessor uncertainty and parsing sensitivity. Matching remained unresolved at the other two targets. These results do not establish that full-vector reconstruction is necessary for internalizing halting, nor that on-axis training is free of instruction-following costs. They show why comparisons of internalization objectives should distinguish achieved shortening, answer correctness and format compliance. The evidence is conditional on selected checkpoints from one model and does not establish training-seed robustness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.