Environment-Aware Large Language Model Routing for Agentic Tasks
Abstract
Large language model routing aims to balance task performance and cost by selecting a suitable model for each query. In this work, we find that the preferred model for an agentic task can change with its execution environment, even when the query remains fixed. Changes in available tools, data scale, and execution feedback affect candidate models' performance, revealing an important source of routing information that query-only routers overlook. We introduce EARoute, an environment-aware routing benchmark of 16k query–environment pairs from seven datasets spanning question answering, coding, and tool use. By varying tools, data scale, interface organization, and execution feedback while holding the query, answer, and evaluation fixed, EARoute shows that environmental changes can reverse model preferences for the same task. We propose a method for making routers environment-aware. Its key challenge is that environments are large and heterogeneous, so they cannot be passed to a router directly and must first be condensed into a compact, informative representation. To this end, we introduce an offline environment-understanding agent that explores environments and compiles routing-relevant properties into reusable extractors. These extractors generate compact descriptions without an LLM call at routing time and can be rerun as environment contents change. Optimized with downstream routing feedback, the resulting representation improves macro-averaged model-success prediction AUC by 8.9 points over query-only inputs, with consistent gains across routing paradigms and held-out environments. Our benchmark and agent-compiled representations demonstrate why execution environments matter for model routing and how to incorporate them effectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.