acceptodds
Under review as a conference paper at ICLR 2027

Tesserae: Scalable Migration Policies for Deep Learning Workloads

Abstract

Training deep learning (DL) models has become a dominant workload in data-centers. Existing schedulers allocate computational resources to running jobs with the goal of optimizing objectives such as makespan, fairness and average job completion time. In addition, they often incorporate placement policies that determine where jobs are assigned within the cluster to improve overall resource utilization. However, as model sizes continue to grow, migration costs have become a major bottleneck for cluster schedulers. Our key insight is that migration problems during scheduling can be formulated as graph-matching problems. Building on this observation, we design novel migration policies that can minimize the number of job migrations during scheduling. We integrate these policies into existing schedulers and demonstrate how our design enables a scalable and efficient GPU cluster scheduler. Experimental results show that Tesserae improves average JCT by up to and reduces the number of migrations by up to compared to existing approaches.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.