Responses Reveal What Outcomes Hide: Response-Informed Value Learning for Multi-Turn LLM Routing
Abstract
Multi-turn language-model routing must select each model before its response is available, although training trajectories retain that response and reveal task success only after the interaction ends. This information timing leaves each intermediate choice without an independent task-quality label while exposing post-call evidence about the return that follows it. We introduce RIVET (Response-Informed Value Estimation for Turn-Level Routing), which jointly trains two value estimators over a shared state–model representation: a deployable Routing Value Estimator (RVE) that scores candidate models from pre-call information, and a training-only Response-Informed Value Estimator (RIVE) that additionally observes the realized response. Both estimators estimate the same discounted return under different information conditions, regress to the same scalar return targets, and are coupled through a shared state–model representation, while inference uses RVE alone. Conditioning on the response turns response-induced return variation from residual noise into an observed covariate, and we show that this cleaner auxiliary regression problem provably reduces the estimation error of the shared representation used by the deployed pre-call router. On 141 held-out MINT tasks with six candidate models, RIVET solves 124 tasks and uses 74.7% less total call cost than MTRouter, which solves 120. The resulting system reaches a high-accuracy, lower-cost operating point without adding a post-call routing path at deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.