ReasonFork: Routing Small and Large Models by Future Consequence
Abstract
Collaboration between small and large language models can make reasoning more efficient by letting a small draft model handle routine steps and consulting a larger teacher selectively. The central challenge is to identify reasoning forks where replacing a draft action with a teacher action can improve task success within a limited step budget while controlling teacher costs. Existing routers based on draft confidence or process scores evaluate the draft's proposed action, but do not explicitly estimate how replacing it with a teacher action would affect subsequent task completion. To address this limitation, we introduce ReasonFork, a lightweight router that learns when to intervene by comparing the future consequences of draft and teacher actions. During training, the draft and teacher each propose an action from the same decision state. An offline evaluator compares the two successors under the same remaining step budget, and the difference in task outcomes provides supervision for the router. At inference, ReasonFork estimates the benefit of replacement using only information available before the teacher call and weighs it against the cost of calling the teacher. At each intervention, the teacher provides one action, after which the draft resumes the task. Experiments across five benchmarks spanning symbolic reasoning and interactive tasks demonstrate the effectiveness of ReasonFork. On reasoning tasks, the reported ReasonFork configurations match or exceed the accuracy of Always Teacher with approximately 72% lower teacher cost, and achieve absolute accuracy gains of 5.4–6.4% over strong baselines at comparable teacher cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.