acceptodds
Under review as a conference paper at ICLR 2027

TROPOS: Risk-aware Scheduling for Distributed Training

Abstract

Computation–communication overlap schedulers rank candidate execution plans under point-estimated operation durations, then run the chosen plan for thousands of iterations in which durations vary with contention. Because durations add along each dependency path and iteration time is the maximum over competing paths, the objective is convex in the durations, and its point-estimate value under-predicts the expected iteration time by a Jensen gap that grows with the number of near-critical paths and with the tail weight of the durations. We formulate overlap scheduling as minimizing the conditional value-at-risk (CVaR) of a stochastic max-plus completion time under memory constraints and show that the deterministic objectives of RoundPipe, LAER-MoE, FCP, HCMS, and NEST are its point-mass special case. TROPOS makes this objective trainable with a log-sum-exp relaxation of the max-plus recursion whose error is bounded and one-sided, a contention-aware duration model with a log-normal–Pareto residual fitted from traces, and tail-focused importance sampling; a heterogeneous graph attention network trained on the resulting loss emits a schedule in one forward pass of 1.24 ms. On Qwen3 dense models from 1.7B to 32B parameters on eight RTX 4090 GPUs, TROPOS lowers P99 iteration latency by 21.0–35.7% and raises throughput by 14–48% relative to RoundPipe, and by a further 8–17% at P99 relative to RoundPipe given the same fitted duration model. Trained only on dense graphs, it lowers P99 on an unseen Qwen3-235B-A22B mixture-of-experts workload by 13.7% relative to LAER-MoE zero-shot and by 20.5% after 50 adaptation steps. The code is available at https://anonymous.4open.science/r/TROPOS.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.