acceptodds
Under review as a conference paper at ICLR 2027

AdaRouter: Learning to Adapt Plans for Agentic Routing via Multiple Large Language Models

Abstract

Agentic routing coordinates calls to heterogeneous large language models (LLMs) and is promising to balance the task performance and inference cost. However, contrary to human problem solving, existing methods for agentic routing lack the ability of reconsideration with feedback evidence gradually acquired during task execution. In this paper, we propose AdaRouter, an adaptive planning framework incorporating reconsideration with available evidence to bridge the gap between human problem solving and agentic routing by LLMs. Specifically, we design two mechanisms for reconsideration, i.e., REFINE that incorporates the new evidence into the query for the next task and REVISE that changes the infeasible plan by updating the remaining tasks or their execution order, while maintaining a coherent guidance across worker calls. Furthermore, we develop a two-level strategy specifically for optimizing AdaRouter with the REFINE and REVISE mechanisms. We achieve trajectory-level GRPO to optimize the overall planning-and-routing behavior and token-level OPSD to distill task-relevant problem-solving skills into routing policy. Extensive experiments across five domains ranging from mathematics to general knowledge and multi-hop reasoning demonstrate that AdaRouter achieves state-of-the-art performance with substantially reduced worker calls in comparison to the strongest routing baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.