CoServe: Preserving Model Choice Under GPU Memory Pressure
Abstract
Multi-model LLM services often offer more models than GPU memory can hold simultaneously. Selecting one model for a request can trigger loading and eviction, even when an already resident alternative could answer it. We present CoServe, which preserves predicted acceptable model alternatives and coordinates request routing with model residency. For each proposed resident set, CoServe recomputes request assignments and compares estimated batch-service costs, waiting penalties, and activation costs. Our structural analysis shows that committing requests to single model identifiers can discard information needed to select a coverage-optimal resident set. On a single B200 deployment with eleven configured model footprints totaling nominal GPU memory, CoServe increases the fraction of requests completed within 100 seconds by at least 16.1 and 3.4 percentage points over CARROT on a Prism-based backend in two arrival realizations. Mean arrival-to-completion latency decreases by 14–17% in the first realization, while tail latency remains higher in both. Numerical ranges reflect timestamp reconstruction uncertainty, not confidence intervals. These results support preserving model alternatives when coordinating routing and GPU residency, with benefits that depend on the workload and latency objective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.