Loop-PP: Pipeline Parallelism for Looped Transformers
Abstract
Looped Transformers increase effective model depth and capacity by recursively applying shared blocks, but their recurrent execution is poorly matched to conventional pipeline parallelism (PP), which assumes a one-pass feed-forward execution order. When recurrent blocks span multiple pipeline stages, outputs from later stages feed back to earlier ones, disrupting conventional PP's fixed compute and communication orders. Supporting recurrence therefore requires globally rescheduling computation and communication. To this end, we present Loop-PP, the first pipeline-parallel training system designed for Looped Transformers. Its core, the Loop-Aware Pipeline Compiler, first makes repeated logical visits and recurrent dependencies explicit in a scheduling-independent Semantic Graph, then synthesizes a Global Compute Order under these semantic constraints, and finally compiles cross-stage dependencies into Per-Rank Action Streams that overlap communication with computation. Across 8B-60B models, Loop-PP maintains broadly comparable throughput to compute-matched baseline with interleaved 1F1B while reducing memory by up to 32.8%. Under an equal resource budget, it improves the best measured throughput by up to 12.7%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.