Decoupling Verification and Generation: Reasoning Content Governs What Fine-Tuning Teaches
Abstract
Language models are increasingly asked to both generate solutions and verifythem, and a single model is often fine-tuned to learn both at once — on the assumption that judging and solving are one competence that improves together. We show they are separately learned: a single fine-tuning run can move the two abilities in different directions. Training on a bare correctness verdict collapses generation almost entirely (math accuracy 63.8% to 0.0%) while leaving verification near the base level, and supervised fine-tuning on solutions (SFT), by contrast, improves generation while leaving verification weak. They are separable because they are learned from different parts of the signal: verification from the judgment it expresses, generation from the reasoning content it retains. A bare verdict carries only the judgment, so it leaves verification intact but collapses generation; an SFT target carries only the reasoning, so it does the reverse; a critique carries both, which is why it appears to improve both at once — not because critique is a uniquely effective objective, but because it is the one signal that contains both ingredients. Varying this content across model sizes and families, we find that whether verification training preserves generation is governed by the reasoning content of the signal, not its quantity. When generation collapses it does so by disengagement rather than lost skill: the model stops attempting problems, a failure we trace to the single instruction framing shared by all critique examples and confirm causally by mixing in direct-solve instructions. The practical consequence is that a verifier cannot be assumed to remain a capable generator; training one that still solves requires keeping reasoning content in the verification signal, and a bare verdict — the cheapest signal — is the most destructive.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.