acceptodds
Under review as a conference paper at ICLR 2027

HostCCL: Decoupling Collective Progress from GPU Execution

Abstract

GPU-driven collective communication uses streaming multiprocessors (SMs) for protocol progress as well as tensor processing. Peer-readiness waits therefore retain GPU resources even when data processing cannot proceed. We present HostCCL, a collective runtime that moves dependency progression to the CPU and uses copy engines for data movement. A CPU executor tracks tile dependencies and submits only ready reductions to a GPU worker, allowing transfer concurrency and reduction parallelism to be configured independently. Batched submission and ordered completion feedback let completed tiles begin forwarding while later reductions execute. On four and eight PCIe-connected L40S GPUs, HostCCL runs AllGather and AllToAll without steady-state GPU kernels, and ReduceScatter and AllReduce with one worker block per GPU. At the evaluated large FP32 reduction points, HostCCL uses 50-75% fewer thread blocks than default NCCL and improves throughput by up to 23.3%. In a controlled four-GPU experiment pairing 256 MiB AllGather with fixed GEMM work, HostCCL completes the joint work in 8 ms, compared with 14.9-15.2 ms for default NCCL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.