acceptodds
Under review as a conference paper at ICLR 2027

Online Serving Problem with Costly Small-model Adaptation

Abstract

Edge-cloud collaboration with large language models offers a promising way to balance model quality and deployment cost: large models provide strong capabilities but incur substantial inference costs, while smaller models are more efficient and can potentially be improved through continuous adaptation. However, existing routing approaches typically treat model capabilities as fixed and only decide which model should serve each request. We study an online serving problem where the learner jointly decides whether to fine-tune a small model and how to serve incoming requests, while adaptation incurs an immediate cost but changes future model performance. We characterize the fundamental difficulty of online adaptation through lower bounds and propose an adaptive algorithm combining convergence detection, optimistic adaptation, and confidence-based serving. The resulting regret guarantee is instance-adaptive and nearly matches the lower bounds. Experiments on edge-cloud LLM serving benchmarks demonstrate that the proposed method effectively exploits adaptation when it improves the small model while avoiding unnecessary fine-tuning when adaptation provides limited benefits.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.