ThinkBranch-OPD: Learning Branch Choices through Counterfactual Distillation
Abstract
A reasoning trajectory often takes shape at a few consequential choices. Once a model selects a substitution or a case split, later steps unfold along that direction. On-policy distillation lets a teacher supervise trajectories actually generated by the student, but feedback is usually delivered token by token along the path already taken. Even when the teacher favors another line of reasoning, that preference is difficult to express as a choice between paths. We introduce ThinkBranch-OPD, which compares the student's realized direction with a teacher-preferred alternative from their shared prefix at points of marked teacher-student disagreement. Supervision thus reaches the choice that determines how the reasoning path continues. Fixed-decision diagnostics show that, after training, the student becomes more inclined toward the teacher-preferred direction at key branches; overall performance on complete solutions also improves. Transferring reasoning ability therefore involves not only learning how to continue a path already begun, but also learning to choose a more promising direction when the path forks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.