FluidPipe: Local Learning for LLM Pretraining and Its Limits
Abstract
Pipeline parallelism (PP) distributes groups of model layers, called stages, across accelerators. The accelerators can sit idle while waiting for gradients from later stages. Local learning removes this wait by training intermediate stages with their own auxiliary losses, but the resulting features may be poorly suited to later stages. We present FluidPipe (FP), a local-learning method for LLM pretraining that adds one decoder block to each intermediate stage’s next-token predictor. These blocks are used only for local prediction and discarded at inference. Throughput gains over PP range from 13 to 73% across the tested configurations from 1B to 70B parameters. At 8B, two-stage FP reaches a nine-task mean accuracy within 0.16 points of full backpropagation after about 15B tokens, with 37% higher throughput on eight H100 GPUs. Using more stages, however, leaves fewer layers in each stage. At 1B, FP and three other local methods suffer an accuracy collapse with four stages of four layers. Across different four-stage partitions of the same sixteen layers, FP’s accuracy ranges from 34.48 to 49.52. Preventing this collapse is the next step toward using local learning’s throughput advantage to reduce LLM pretraining cost across more accelerators.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.