Learned subspace compression for communication efficient pipeline parallelism
Abstract
Pipeline parallelism enables training of large language models (LLMs) that exceed single-device memory but requires frequent communication of activations between pipeline stages. On bandwidth-constrained networks, this communication overhead can become the dominant training bottleneck in decentralized environments. The existing approach, Subspace Networks, mitigates this by constraining activations and model weights to a fixed, shared low-rank subspace across all layers. However, this global constraint restricts the model's representational capacity and leads to substantial performance degradation when trained on the same number of tokens. We propose Manifold Aware Projection Learning (MAPL) for communication efficient pipeline parallelism, a framework that formulates inter-stage activation compression as learnable orthogonal projections optimized directly on the Stiefel manifold. Rather than imposing a fixed global subspace, MAPL enables each pipeline stage to dynamically adapt its own compression subspace, coupled with factorized anchor embeddings that restore token-specific offsets with negligible bandwidth overhead. Across LLaMA models from 150M to 3B parameters, MAPL achieves 4 to 11 activation compression with minimal degradation in validation loss relative to uncompressed training. % On the 3B model, MAPL narrows the validation loss gap relative to uncompressed training from 18% (with Subspace Networks) down to just 1.6%. Finally, in a more challenging regime under emulated 1Gbps bandwidth links, MAPL scales to an 8B-parameter model at an extreme 102 activation compression (85× end-to-end) ratio while achieving faster convergence than uncompressed training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.