Autonomy and Performance Trade-off in Adaptive LLM Cascades
Abstract
We present an adaptive system where two models, a large teacher and a small student, form a cascade in an online deployment scenario. As the queries stream in, a router decides whether the teacher model is needed for each query. We continually update the router and distill from the stronger, larger model to the smaller student, whenever the larger model is queried. This creates a non-stationary system, with both the router and the student constantly adapting. The process trades-off the overall performance of the system and the autonomy of the small model, which reflects the privacy of the computation. We treat routing as a contextual bandit problem, and experiment with several non-stationary data stream distributions. Experiments demonstrate the effectiveness of our method, and show that competing incentives can be balanced to significantly advance the Pareto efficiency of LLM serving, even without access to the teacher’s logits (i.e., with API models).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.