Propose-and-Refine: Multi-Turn LLM Routing for Long-Horizon Tasks with Virtual-Rollout MCTS
Abstract
Large language models (LLMs) are increasingly deployed for long-horizon tasks, where inference costs accumulate over multiple turns and different turns may require different model capabilities. Unlike episode-level routing, multi-turn routing selects a model at each interaction step to better balance performance and cost under user preferences. However, learning effective multi-turn routers remains challenging because the long-term value of a routing action depends on future routing plans, and high-quality supervision is difficult to obtain under a limited real-environment rollout budget. We propose ProRe, a propose-and-refine multi-turn routing framework. At each state, a generative proposer samples routing plans composed of variable-duration macro-actions, and a plan-conditioned refiner ranks them by predicted terminal performance and cumulative cost under a user-specified preference, thereby incorporating future information into the current decision. To train ProRe efficiently, we introduce OnVir, an online data-collection strategy that replaces real-environment rollouts in Monte Carlo tree search (MCTS) with ProRe-generated virtual rollouts, discovering higher-quality routing trajectories under a fixed sampling budget. Extensive experiments show that ProRe and OnVir jointly improve routing performance across diverse long-horizon tasks and user preferences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.