TopoGAT: Topology-Aware Unsupervised Root-Cause Localization in Distributed Training via Spatio-Temporal Graphs
Abstract
Large language model (LLM) training relies on parallelism for efficiency, yet this architecture introduces a vulnerability: a single faulty rank propagates delays through collective synchronization barriers, slowing the entire cluster and causing substantial computational waste. Existing fault localization methods depend on fault labels or hand-crafted rules, which are scarce in production and generalize poorly across new training workloads. They also treat ranks independently, ignoring the training topology and failing to distinguish root-cause ranks from victim ranks slowed by synchronization. We propose TopoGAT, an unsupervised spatio-temporal graph framework for root-cause rank localization. TopoGAT exploits the inherent parallelism topology by reconstructing each rank’s behavioral profile from its topological neighbors and its own history, learning normal synchronization patterns without fault labels. Using universally available telemetry—collective communication events, GPU metrics, and network counters—it generalizes across diverse workloads. Trained once on a single benchmark topology and evaluated zero-shot on 11 real-world datasets spanning 3 model families, 10 parallelism configurations, and both computation- and communication-side faults, TopoGAT achieves an average top-3 localization accuracy (Hit@3) of 94.6%, significantly outperforming state-of-the-art baseline methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.