acceptodds
Under review as a conference paper at ICLR 2027

DBO++: Rethinking Communication Overlap for Dense Tensor-Parallel Prefill

Abstract

Communication overlap can reduce prefill latency in dense language models, but its value depends on more than the input batch. We present DBO++, which extends vLLM's dual-batch execution to dense tensor-parallel prefill, and study how engine configuration changes the choice between overlapping and unsplit execution. On Qwen3-32B with four H200 GPUs, changing the configured batch-token limit while keeping the input and split fixed changes the mean effect from a 10.1% slowdown to a 3.6% latency reduction. Controlled comparisons connect this reversal to startup profiling and memory allocation. In the measured setting, live allocation increases allocator retries and imposes a penalty concentrated in the overlapping path. A targeted intervention preserves existing KV capacity while changing subsequent allocation handling; measured retries disappear and a 24.6% live-allocation penalty falls to approximately zero. Separately, a matched comparison with fixed KV allocation establishes a 2.1% latency reduction over unsplit execution. A rule using workload information and the configured limit preserves a token-only rule's positive savings while avoiding its two harmful choices in an initial held-out evaluation. Independent follow-up tests expose output mismatches and slower selections, delimiting this empirical rule's support. A broader four-model, one-output-token survey records both gains and regressions. Together, these results establish configuration-dependent overlap in the studied system and show why execution selection must be evaluated with the engine conditions and output contract under which it is used.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.