LEROS: Learning–Evolution Reinforced Orchestration for Dynamic LLM Workloads on Cloud-Native Kubernetes Clusters
Abstract
Large language model (LLM) clusters now run pre-training, fine-tuning and online serving side by side on one Kubernetes fleet, yet the two workload classes are still orchestrated by separate, reactive controllers: gang schedulers hold GPUs until training finishes, while horizontal pod autoscalers scale serving on lagged utilisation. Production traces from a 1024-GPU fleet show this split design reaches only 31.2% mean GPU utilisation and causes 16.8% serving SLO violations during flash crowds. We present LEROS, a two-layer framework that unifies training and serving under one forward-looking controller. The offline layer uses NSGA-II over a seven-parameter policy genome on a calibrated digital twin; the Pareto knee and all archived individuals are validated on hardware replay (Spearman , ). The online layer is a GNN actor–critic trained with PPO, warm-started by behaviour cloning from the knee point and continuously fine-tuned; a drift detector re-triggers the search when workloads shift. On a 128-node/1024-GPU production cluster (A100+H100, InfiniBand NDR 400 Gbps), via trace-driven replay with real PyTorch/Megatron-LM jobs and real serving load, LEROS raises utilisation 31.2%76.4%, cuts SLO violations 16.8%2.9%, reduces makespan 48.6 h26.7 h (45.2% vs default, 8.0% vs the strongest separated baseline; paired bootstrap , ), and lowers energy 27.625.6 MWh by powering down 18 of 128 nodes. Ablations attribute a 3.3% stationary-workload makespan gain to the RL head and a 12.8% gain under drift. All code, manifests and anonymised traces are released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.