A Theoretical Study of Model Routing with Distributional Losses
Abstract
Generative AI systems increasingly rely on routing user prompts among multiple generative models with different capabilities, costs, and output qualities. This is the problem of model routing, where a decision-maker selects among different generative models in an online fashion to optimize the quality and minimize the cost of the generated corpus. Currently, the majority of model routing algorithms use hard-coded routing choices or assume that the generated output can be evaluated in a per-prompt manner. However, a per-prompt loss alone cannot capture the efficacy of a generative model; instead, one needs to evaluate the generated output at a *distributional* level. For example, the Fréchet Inception Distance (FID) depends on the distribution of generated images. Motivated by this example, we introduce a formulation of the online model routing problem where the loss can depend on both per-round and distributional-level quantities. We show that one can extend classical decision-making algorithms such as Upper Confidence Bound (UCB) and Explore-Then-Commit to handle such losses when the loss is convex and satisfies a tail-bound assumption. These conditions are satisfied by canonical examples used widely in practice such as FID and Frontier Integral (FI). Overall, our work develops a principled formulation of model routing with novel distributional losses and shows that principled decision making is possible in these challenging settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.