Memory-Aware Routing for Cost-Efficient Language Agents
Abstract
Language-model agents must balance the reliability of strong models against the cost of invoking them at every decision step. Retrieved experience offers another way to manage this trade-off: a past trajectory may provide the support a smaller model needs to match a stronger reference. However, memory can also mislead, and its usefulness depends on both the current state and the model that consumes it. Routing therefore requires deciding not only which model to invoke, but also which memory can make that model sufficient. We formulate this problem as cost-constrained selection over memory-model pairs and develop a two-stage framework that separates relative memory benefit from absolute routing reliability. The first stage shortlists memories by predicted alignment gain; the second uses historical outcome evidence to estimate model-specific reliability, with a shared-bias-low-rank formulation for cross-model prediction. We evaluate an evidence-grounded approximation on OfficeBench, ScienceWorld, and -bench, alongside separate component studies of sharing and shortlisting. The evaluated router achieves up to 17.8% proxy cost savings over always invoking the reference, routing 25.7% of states to smaller models at 82.0% routed precision. These findings support treating episodic memory can effectively serve as part of model allocation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.