PipeTempo: Stagewise Accumulation for Asynchronous Pipeline-Parallel Training
Abstract
In GPU-rich language-model pretraining, large synchronous batches can improve hardware utilization but leave fewer optimizer updates within a fixed token budget. We show that more frequent asynchronous updates can improve model quality in this regime, despite using stale gradients. We introduce PipeTempo, which sets each pipeline stage's accumulation interval from its local delay. Later stages use smaller batches and update more often, enabling more frequent updates than PipeDream-2BW while limiting local staleness to one update and retaining at most two weight versions per stage. We generalize the convergence analysis of stochastic conditional-gradient updates with momentum to coupled pipeline stages with arbitrary fixed accumulation intervals. Under an expected -Kurdyka–Łojasiewicz condition, we prove convergence for both PipeTempo and PipeDream, preserving the synchronous method's leading stochastic complexity for expected loss gap . Language-model pretraining experiments with Muon on FineWeb-Edu show lower validation loss than synchronous training, PipeDream, and PipeDream-2BW at matched token budgets in the large-batch regime.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.