acceptodds
Under review as a conference paper at ICLR 2027

Block Parallelism for Efficient Distributed Long-Context Diffusion Language Model Training

Abstract

Block diffusion language models (BDLMs) combine autoregressive dependencies across blocks with parallel denoising within blocks, but long-context training is constrained by distributed attention communication and activation memory. Conventional context parallelism (CP) shards the combined clean-plus-corrupted sequence by position, communicating shared clean K/Vs together with block-specific corrupted K/Vs and their gradients. Our key observation is that the BDLM objective separates over target blocks and that diffusion-based speculative drafter losses have the same decomposition. We introduce block parallelism (BP), a new distributed parallelism dimension that assigns each corrupted-block computation to one rank. To scale BP to long contexts, we introduce context-sharded block parallelism (CSBP), which also shards the shared clean sequence across those ranks. CSBP keeps corrupted K/Vs and their gradients local, avoids replicated clean prefixes, and preserves BDLM training semantics. Applied to speculative drafters, it also avoids replicated block-local computation. On 16 H200 GPUs at 256K context, CSBP improves throughput over the best baseline by \textbf{1.18--1.45\times} for supervised fine-tuning on 3B–26B-A4B models and \textbf{1.27--1.33\times} for converting 27B autoregressive models to BDLMs, while matching or reducing peak HBM. For DiffusionGemma 26B-A4B, the fine-tuning speedup over the best baseline grows with context length, reaching \textbf{1.61\times} at 512K. On eight H100 GPUs, CSBP accelerates Qwen3.8-27B DFlash2 drafter training by \textbf{2.48\times} at 512K and \textbf{7.59\times} at 1M. In 12-hour DiffusionGemma 26B-A4B SFT, CSBP achieves higher pass rates at every trained checkpoint on SWE-bench Verified and Terminal-Bench Lite.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.