acceptodds
Under review as a conference paper at ICLR 2027

Autoscaling by Imagination: Graph World Models for Safe Distributed ML Services

Abstract

Modern machine learning services are increasingly deployed as distributed microservice systems, where resource allocation decisions must balance latency, reliability, and cost. In practice, however, autoscaling remains largely reactive: controllers add replicas only after the system is already under pressure, while safer deployments often rely on static over-allocation of resources. Existing autoscalers do not learn a model of how scaling actions propagate through the service graph, and therefore cannot anticipate how a decision will affect future load, latency, cost, and service-level violations before it is executed. To fill this research gap, we present NetDream, a graph world model for safe autoscaling of distributed machine learning services. NetDream learns action-conditioned dynamics over a microservice graph, predicting how load, latency, replicas, and violation risk evolve under candidate scaling actions. At runtime, NetDream uses this learned model for imagination-based planning: it rolls out multiple future scaling trajectories, filters out plans predicted to violate service-level constraints, and executes the first action of the best safe plan. This turns autoscaling from threshold-based reaction into topology-aware planning over a learned model of the system. We evaluate NetDream on a production-grade AWS Kubernetes deployment of Online Boutique, an 11-service, 14-edge microservice graph with a 66-dimensional system state and 48,000 collected transitions, under steady, variable, bursty, and flash-crowd traffic. NetDream is the most cost-efficient safe autoscaler among all evaluated methods. Each marginal replica-step it spends buys safety 2.7 times more costefficiently than the best static over-provisioning strategy. On variable-traffic workloads specifically, it achieves about 5 times fewer violations than reactive autoscaling while using only about half the cost of the strongest over-provisioning baseline. Code, trained models, and infrastructure manifests are released anonymously at https://anonymous.4open.science/r/netdream-FCF9.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.