acceptodds
Under review as a conference paper at ICLR 2027

Asynchronous Large-to-Small Handoff for Efficient Reasoning

Abstract

Long chain-of-thought (CoT) reasoning improves the performance of large language models on complex tasks, but also incurs substantial inference cost and latency. We investigate whether a smaller model can help complete an ongoing large-model reasoning trajectory faster while preserving the larger model's answer quality. We propose , a framework that launches small-model continuations from intermediate large-model reasoning states, creating an online race between continued large-model reasoning and candidate small-model completions. We further train a prefix-conditioned Reasoning Finisher with reinforcement learning to improve the accuracy–latency trade-off. Across ten benchmarks spanning mathematics, science, and code generation, ALSH reduces latency by compared with standalone large-model inference in our main Qwen3.5-35B/Qwen3.5-2B setting, with a relative improvement in dataset-macro accuracy, demonstrating a stronger accuracy–latency balance than prior large–small collaboration methods. Experiments across different model scales and families further support the generality of ALSH.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.