Right Answer, Wrong Reasoning in Verifier-Filtered Self-Training
Abstract
Self-training for reasoning keeps a sampled solution only if its final answer is verified correct, and recovers problems the model cannot solve by showing it the answer. In logic and planning puzzles where every intermediate claim can be checked exactly, a verified answer turns out to say little about the reasoning behind it. The claims in a solution are of two kinds: forced claims, which the answer depends on, such as the moves of a plan, and unforced claims, which it does not, such as descriptions of the state along the way. Any solution the verifier accepts has its forced claims right almost by construction; its unforced claims go unchecked, and in accepted solutions written with an answer hint up to 74% of them are false. On Knights-and-Knaves problems the model could not solve, 67.1% of these solutions contain a false claim; giving the same information as the opening steps of a solution lowers this to 0.4%. The false claims do not stay in the data. Models trained on accepted but unsound solutions give correct answers resting on a false claim 53–75 percentage points more often than models trained on sound solutions to the same problems, with no consistent difference in accuracy, and the same inheritance appears in planning. Filtering the accepted solutions with the checker before training removes almost all of it. The effect appears in all three model families measured (Qwen, Llama, and OLMo) and, in Qwen3 and OLMo-3, shrinks with scale without vanishing. An outcome verifier certifies answers, not reasoning, and a model trained on what it certifies can learn to be right for the wrong reasons.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.