ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism
Abstract
Unlocking large-scale, low-bandwidth decentralized training has the potential to utilize otherwise untapped compute resources. In centralized settings, large-scale multi-node training is primarily enabled by data and pipeline parallelism, two techniques that typically rely on high-bandwidth communication. While efficient methods now exist for decentralized data parallelism, pipeline parallelism remains a major challenge. Recent efforts, such as Subspace Models (SM), report up to activation compression but rely on constrained optimization of output projection layers. We propose a different approach: the Residual Bottleneck Model (ResBM), an architecture designed from the ground up for low-bandwidth communication environments while remaining applicable to standard transformer-based architectures. ResBM introduces a residual encoder–decoder bottleneck module across pipeline boundaries that can be trained end-to-end with standard optimizers and preserves an explicit low-rank identity path. In our main comparative experiments with up to 3.5B-parameter models trained for up to B tokens on large-scale web datasets, we show that ResBM achieves the lowest loss among evaluated methods at extreme compression rates of and . Although the authors of SM have reported little or no degradation from activation compression, our longer-horizon, optimizer-matched evaluations reveal a measurable loss gap for every compression method evaluated, including ResBM. However, ResBM consistently exhibits the smallest such gap relative to uncompressed baselines. Its loss gap remains approximately constant during later training, whereas SM's appears to grow over the evaluated horizon. We therefore position ResBM as a new state of the art for extreme activation compression in low-bandwidth pipeline-parallel training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.