SSN++: Layerwise Subspace Networks for Low-Bandwidth Pipeline Training
Abstract
Subspace networks (SSNs) make pipeline-parallel training communication-efficient by constraining every layer to write into one shared low-dimensional subspace, so activations and gradients cross each stage boundary losslessly. This shared subspace, however, limits what the model can express, and with it the loss an SSN can reach relative to an uncompressed transformer. We propose SSN++, which removes this constraint without additional communication. We introduce a norm-preserving orthogonal transform that losslessly transports each compressed activation from one layer's subspace into the next, so every layer can write into its own subspace. We extend this separation within each layer, giving attention and the MLP separate subspaces and dividing each into groups with their own subspaces, which together exceed the rank of one subspace at equal parameters, and we add learned gates that choose what each layer passes on. Because the compressed activations of preceding tokens have already crossed the boundary, the receiving layer also decodes its input from them with a learned causal filter, raising its rank beyond one subspace. Random subspaces spread activations like rotation-based quantisers, which makes the communication amenable to quantisation, and we spend the saved bits on a higher rank. Approximation theory and measurements during training support each choice. From 229M to 2B parameters, SSN++ lowers loss and raises eval accuracy over a parameter-matched SSN at equal communication, and with quantised communication it closes 63 to 78% of the matched SSN's loss gap to the uncompressed transformer at 229M and 1B and 58% at 2B, while training 12 to 26 times faster per update than it over 200 Mb/s links. The loss gap of SSNs thus stems largely from how capacity is organised around the compressed activations, and closing it without extra communication makes pipeline-parallel training over internet-grade links far more efficient.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.