acceptodds
Under review as a conference paper at ICLR 2027

Learning Teacher Output Preferences from Final Responses for Reasoning Distillation

Abstract

Large reasoning models often expose only their final responses while keeping their native chain-of-thought (CoT) unavailable, limiting their use as teachers for smaller models. Training directly on final responses omits intermediate supervision, while generic synthetic rationales may overlook source-dependent output patterns. We investigate whether visible final responses provide signals for synthesizing reasoning supervision that reflects these patterns. We introduce Final-Output Preference Learning (FOPL), a final-response-conditioned preference-learning framework. In the first offline stage, we use auxiliary trace-accessible source teachers to pair each native CoT with answer-conditioned synthetic alternatives generated by strong reasoning models. Each preference pair shares the same question and visible final response, allowing a single reconstruction policy to learn preferences over reasoning realizations while holding the response fixed. At the second transfer stage, the policy generates a CoT from only a question and the target teacher's visible final response, without access to that teacher's identity or native CoT. Reconstructions conditioned on a teacher family's own responses resemble its native traces more closely than those conditioned on other families' responses, including for a family excluded from preference training. On several completed reasoning evaluations, students trained on reconstructed CoTs match or outperform students trained on native-CoT references. These results suggest that visible final responses can support useful reasoning supervision without requiring the recovery of a target teacher's hidden reasoning process.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.