D-PDLP: Scaling First-Order Primal–Dual Optimization to Distributed Multi-GPU Systems
Abstract
We present a distributed framework of the Primal-Dual Hybrid Gradient (PDHG) algorithm for solving massive-scale linear programming (LP) problems. Although PDHG-based solvers perform strongly on single GPUs, computational throughput can become a performance bottleneck for sufficiently high-workload instances. To overcome this bottleneck, we propose D-PDLP, a distributed PDLP framework and, to our knowledge, the first to accelerate the complete PDLP solver across multiple GPUs. By combining two-dimensional partitioning with our load-balancing strategy, D-PDLP distributes primal-dual computation across GPUs while maintaining efficient GPU execution. To understand how D-PDLP scales across GPUs, we derive a PDLP-specific cost model that relates computation and communication costs to the processor-grid configuration and GPU count. The model guides resource selection, with case studies on contrasting sparse workloads demonstrating its predictive value. Extensive experiments on standard LP benchmarks, including MIPLIB and Mittelmann instances, and huge-scale real-world datasets show that our implementation, built upon cuPDLPx, achieves approximately speedup on eight GPUs for computationally intensive instances while maintaining numerical accuracy. Our solver is released as open-source software.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.