Cloud on the Clock: Online LLM Routing with Early Edge Takeover
Abstract
An edge–cloud LLM system routes each query to either a local edge model or a cloud model to balance answer quality against user waiting time. In practice, many routers are trained offline and do not adapt to new service feedback. Some recent routers learn model choices online, but cloud exploration can incur prolonged waiting and highly variable utility feedback, making exploration less efficient. The deadline-only service runs the selected model until delivery or a final deadline, which cannot resolve the difficulty of exploring cloud and learning the quality-delay trade-off under such variability. Therefore, we propose an online edge-cloud routing framework that combines early edge takeover with kernel upper confidence bound (UCB) routing based on semi-adaptive (SA) confidence calibration. During service, the framework interrupts a cloud request that remains unfinished at a threshold and starts an edge generation attempt. After each round, terminal quality–delay feedback updates the selected initial route's weighted kernel utility estimate, while the corresponding prediction error calibrates the SA confidence bonus. Both updates guide UCB routing for the next query. When cloud delays exhibit heavy tails, we analyze the framework using a conditional Pareto model motivated by Pareto execution-time models in cloud straggler analysis. Our analysis establishes conditions for vanishing average regret and for a strictly smaller variance-dependent cumulative regret bound than deadline-only service. Comparing with the strongest baselines on six benchmarks, our method achieves a mean relative utility gain of 27.9% and reduces p95 latency by 36.5%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.