acceptodds
Under review as a conference paper at ICLR 2027

Not Every Step Matters: Learning Multi-Turn LLM Routers

Abstract

LLM agents complete tasks through a sequence of model calls and environment interactions, causing inference costs to accumulate. Routing each call to an appropriate model offers a way to reduce these costs while maintaining task performance. However, the capabilities needed can change over the course of a task, and choosing a different model does not matter equally at every step. We find that the value of model selection is concentrated at a small subset of steps. We call these steps pivotal: differences in the capabilities of candidate models at the same agent history lead to substantially different final task outcomes. A step can be difficult or critical to the task without being pivotal if changing the model makes little difference. We identify pivotal steps through counterfactual rollout trees and use their model comparisons to guide data collection, router training, and inference. Our router matches or outperforms the strongest fixed-model baseline on five of six interactive benchmarks. It also reduces total inference cost by 47.8% across the 473-task evaluation suite.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.