acceptodds
Under review as a conference paper at ICLR 2027

Kernel-Level Context Parallelism: Bridging Incompatible Shard Layouts in Hybrid Linear-Attention Transformers

Abstract

Hybrid transformers that interleave softmax attention with linear attention now reach production scale, but under context parallelism (CP) their two layer families require incompatible sequence layouts. Fused softmax-attention CP kernels shard each sequence in a zigzag order to balance causal work, whereas state-passing linear-attention kernels require each rank to own a contiguous span. A common workaround all-gathers the full hidden state at every linear-attention layer and recomputes the layer redundantly on all ranks, so its activations and linear-attention compute stay proportional to the full sequence however large grows. At 256K tokens with , this baseline exhausts an 80 GB H100 before completing a single step. We present KCP, which instead converts between the two layouts with an exactly invertible token permutation. KCP keeps zigzag as the global layout, converts to contiguous order at the entry of each linear-attention layer with a single all-to-all of tokens per rank, runs the unmodified native CP kernel, and converts back on exit. The payload of each collective drops by ; more importantly, per-rank linear-attention compute and activation memory also drop by , because no rank recomputes the full sequence. Because the conversion is a pure permutation, its backward pass is the reverse permutation. On a 35B-parameter MoE hybrid across 32 H100 GPUs, KCP achieves 2.5–4.7× the training throughput of the all-gather baseline at every configuration both can run, and it trains the 256K-token configuration at 53.3 GB peak allocated memory. Training losses agree with the baseline up to floating-point nondeterminism, and the layout conversion accounts for 1.8–2.7% of GPU-busy time across profiled ranks. On a dense hybrid with small bins, where the gather is not the bottleneck, the two methods are within ±7% at three of four points, consistent with gains that come from eliminating full-sequence materialization rather than from a faster kernel.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.