Internalizing the Halt Vector: Full-Vector and On-Axis Reconstruction
Abstract
Can a reasoning model learn earlier stopping by matching only the projection of a steered activation, rather than reconstructing the full vector? We compare full-vector and on-axis reconstruction on DeepSeek-R1-Distill-Qwen-7B, separating reasoning length, mathematical answer correctness and compliance with the requested answer format. An outcome-informed follow-up to an earlier held-out comparison, with development-only selection, evaluated three shortening targets on 2,000 held-out problems. One target established matched shortening within one percentage point: full-vector and on-axis reconstruction reduced thinking by 23.56% and 23.39%, respectively. Their accuracy was equivalent within the prespecified 2.6-point margin under Math-Verify, while a boxed-answer grader favored full-vector reconstruction by 2.80 points. The on-axis model produced more closed, unboxed answers despite an explicit boxing instruction. A post-hoc, model-assessed audit supports a predominantly formatting-based explanation of this discrepancy, subject to assessor uncertainty and parsing sensitivity. Matching remained unresolved at the other two targets. These results do not establish that full-vector reconstruction is necessary for internalizing halting, nor that on-axis training is free of instruction-following costs. They show why comparisons of internalization objectives should distinguish achieved shortening, answer correctness and format compliance. The evidence is conditional on selected checkpoints from one model and does not establish training-seed robustness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.