acceptodds
Under review as a conference paper at ICLR 2027

TopoGAT: Topology-Aware Unsupervised Root-Cause Localization in Distributed Training via Spatio-Temporal Graphs

Abstract

Large language model (LLM) training relies on parallelism for efficiency, yet this architecture introduces a vulnerability: a single faulty rank propagates delays through collective synchronization barriers, slowing the entire cluster and causing substantial computational waste. Existing fault localization methods depend on fault labels or hand-crafted rules, which are scarce in production and generalize poorly across new training workloads. They also treat ranks independently, ignoring the training topology and failing to distinguish root-cause ranks from victim ranks slowed by synchronization. We propose TopoGAT, an unsupervised spatio-temporal graph framework for root-cause rank localization. TopoGAT exploits the inherent parallelism topology by reconstructing each rank’s behavioral profile from its topological neighbors and its own history, learning normal synchronization patterns without fault labels. Using universally available telemetry—collective communication events, GPU metrics, and network counters—it generalizes across diverse workloads. Trained once on a single benchmark topology and evaluated zero-shot on 11 real-world datasets spanning 3 model families, 10 parallelism configurations, and both computation- and communication-side faults, TopoGAT achieves an average top-3 localization accuracy (Hit@3) of 94.6%, significantly outperforming state-of-the-art baseline methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.