acceptodds
Under review as a conference paper at ICLR 2027

Responses Reveal What Outcomes Hide: Response-Informed Value Learning for Multi-Turn LLM Routing

Abstract

Multi-turn language-model routing must select each model before its response is available, although training trajectories retain that response and reveal task success only after the interaction ends. This information timing leaves each intermediate choice without an independent task-quality label while exposing post-call evidence about the return that follows it. We introduce RIVET (Response-Informed Value Estimation for Turn-Level Routing), which jointly trains two value estimators over a shared state–model representation: a deployable Routing Value Estimator (RVE) that scores candidate models from pre-call information, and a training-only Response-Informed Value Estimator (RIVE) that additionally observes the realized response. Both estimators estimate the same discounted return under different information conditions, regress to the same scalar return targets, and are coupled through a shared state–model representation, while inference uses RVE alone. Conditioning on the response turns response-induced return variation from residual noise into an observed covariate, and we show that this cleaner auxiliary regression problem provably reduces the estimation error of the shared representation used by the deployed pre-call router. On 141 held-out MINT tasks with six candidate models, RIVET solves 124 tasks and uses 74.7% less total call cost than MTRouter, which solves 120. The resulting system reaches a high-accuracy, lower-cost operating point without adding a post-call routing path at deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.