HIDDEN-STATE GEOMETRY FOR REASONING SHARPENING VIA LAYER-WISE TUNING
Abstract
A wrong answer can be informative about what a reasoning model still needs to learn. This is especially true when the model assigns almost the same score to a correct answer and an incorrect competitor, yet prefers the latter. We study these small-gap errors as a problem of learning a specific distinction. Candidate diagnostics show close competitions in the inspected pools, with incorrect predictions nearer the support boundary than correct predictions. This suggests using the competing answer to guide correction, rather than supplying the correct response alone. We introduce Layer-Selected Geometry Repair: select a hidden layer on held-out candidate data, build a fixed correct-minus-wrong direction from training pairs, and train a pairwise hidden margin alongside an anchored answer-learning objective. The comparison is internal, but its purpose is practical: improve the returned answer without adding a hidden-space selector at inference. Mathematical and medical evaluations support the procedure, while separate diagnostics examine adaptation, layer choice, and direction. A close wrong answer is not only a failure to discard; it can identify the distinction that training should address.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.