Joint Communication-Computation Optimization For Multi-GPU Heterogeneous GNN Training
Abstract
Full-graph training of heterogeneous graph neural networks (GNNs) remains difficult to scale because irregular relation-specific workloads and cross-GPU aggregation jointly limit computation and communication efficiency. This bottleneck is particularly acute in learning-assisted large-scale supply chain production planning, where candidate-variable screening depends on preserving long-range heterogeneous dependencies, making sampling an unsuitable substitute for fullgraph execution. We propose a phase-aware multi-GPU framework that jointly optimizes workload organization and cross-GPU communication while preserving message-passing semantics. At preparation time, a lightweight surrogate selects the graph workload granularity by trading scheduling flexibility against one-time materialization overhead. During recurrent training, asynchronous aggregation overlaps cross-GPU reduction with local computation, while a net-benefit controller selects the communication granularity and reverts to blocking execution when overlap is not worthwhile. The resulting transformations preserve relation-specific full-graph aggregation up to floating-point reduction order. Experiments on five consecutive industrial supply-chain graphs show that the selected preparation configuration is within 0.709% of the empirical optimum on a held-out snapshot. The complete framework reduces mean steady-state epoch time by 62.50% and 50-epoch wall-clock time by 58.50%, with the additional preparation cost amortized after approximately four epochs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.