acceptodds
Under review as a conference paper at ICLR 2027

DitronGen: Compiler–LLM Coordination for Adaptive Kernel Generation and Computation–Communication Overlap

Abstract

Optimizing training and inference of large artificial intelligence models on distributed platforms is challenging because inter-device communication can limit efficiency of computation. High-level frameworks typically optimize computation and communication separately, while manually implementing overlapped kernels demands substantial expertise. We present DitronGen, a framework that coordinates a compiler and an LLM agent for adaptive kernel generation and computation–communication overlap. Starting from single-device programs with user-specified tensor placements, the compiler infers communication requirements and adapts computation and communication implementations to hardware architecture, interconnect topology, and communication dimensions. The LLM agent generates synchronization code using Triton-distributed primitives, adapting synchronization operations and their placement to the data dependencies and control flow of the generated kernels to enable tile-level overlap across different communication scenarios. This division of responsibilities coordinates compiler-based generation of computation and communication kernels with LLM-based generation confined to synchronization. DitronGen can be used as a custom backend for \tt torch.compile. On NVIDIA Hopper GPUs, it achieves speedups of 1.13×–3.70×, 4.07×–15.3×, and 2.55×–5.59× over Torch Inductor for GEMM, Group GEMM, and Attention, respectively, and up to 1.14× and 5.22× for dense MLP and MoE layers.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.