DitronGen: Compiler–LLM Coordination for Adaptive Kernel Generation and Computation–Communication Overlap
Abstract
Optimizing training and inference of large artificial intelligence models on distributed platforms is challenging because inter-device communication can limit efficiency of computation. High-level frameworks typically optimize computation and communication separately, while manually implementing overlapped kernels demands substantial expertise. We present DitronGen, a framework that coordinates a compiler and an LLM agent for adaptive kernel generation and computation–communication overlap. Starting from single-device programs with user-specified tensor placements, the compiler infers communication requirements and adapts computation and communication implementations to hardware architecture, interconnect topology, and communication dimensions. The LLM agent generates synchronization code using Triton-distributed primitives, adapting synchronization operations and their placement to the data dependencies and control flow of the generated kernels to enable tile-level overlap across different communication scenarios. This division of responsibilities coordinates compiler-based generation of computation and communication kernels with LLM-based generation confined to synchronization. DitronGen can be used as a custom backend for \tt torch.compile. On NVIDIA Hopper GPUs, it achieves speedups of 1.13×–3.70×, 4.07×–15.3×, and 2.55×–5.59× over Torch Inductor for GEMM, Group GEMM, and Attention, respectively, and up to 1.14× and 5.22× for dense MLP and MoE layers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.