acceptodds
Under review as a conference paper at ICLR 2027

KV-Pipe: On the Relation Between KV Sharing and Pipeline Parallel Efficiency in LLMs

Abstract

Pipeline parallelism (PP) is bounded by its slowest stage. In LLMs, however, the stages are not identical: the last stage also carries the LM head, so an evenly split decoder already leaves every other device idle for part of each step. Moving whole layers between stages cannot fix this, because a single layer is coarser than the imbalance itself. We observe that cross-layer KV sharing, which has so far been studied as an inference-time memory optimization, is exactly the fine-grained knob this problem needs. Each shared layer removes a small, predictable amount of compute from its stage. We present , which places shared-KV layers on the bottleneck stage and chooses how many to share by minimizing a FLOPs Imbalance Ratio (FIR) computed offline from per-layer FLOP counts. For LLaMA2-7B with PP8 on Ascend 910B NPUs, sharing KV in only four layers removes 1.5% of model FLOPs but cuts iteration time by 9.8% and raises MFU from 56.3% to 61.5%. Across the sweep, FIR orders measured MFU almost perfectly (Spearman at PP8). Sharing more layers makes both FIR and MFU worse again, even though total FLOPs keep falling. When we vary the LM-head cost, the budget predicted from FIR matches the measured optimum in all five settings. The gains grow with pipeline depth, also in a controlled sweep that fixes everything but the number of stages. On LLaMA2-7B and LLaMA3-8B, validation perplexity and downstream accuracy stay within the run-to-run variation of full attention. The gains are additive to the Seq1F1B schedule, carry over to LLaMA2-13B, LLaMA3-8B and Qwen2.5-14B on GPUs, and the same kind of late-layer sharing speeds up GQA decoding by 7–8%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.