When Is Low-Rank Pretraining Actually Efficient? A Resource-Aware Benchmark
Abstract
Low-rank methods offer a broadly applicable way to improve the efficiency of LLM pretraining by reducing model parameters, optimizer memory, or FLOPs. Yet their practical efficiency remains difficult to compare. Existing studies use different hardware, model scales, implementations, and training budgets, and often report intermediate savings without showing how they translate into end-to-end memory and training time. Therefore, a meaningful comparison requires a shared training setup, explicit control of the training budget, and direct measurement of both model quality and realized system costs. To provide such a comparison, we introduce the Low-rank Efficient Pretraining Benchmark (LR-EPB), which evaluates ten low-rank pretraining methods that place low-rank structure in model weights, optimizer states, or weight updates, on LLaMA-style models from 60M to 1B parameters. Across 130 runs and approximately 13,000 equivalent A100-40 GPU-hours, we compare model quality under matched-token and matched-FLOPs budgets and measure training FLOPs, peak memory, throughput, and serving latency. Our results show that low-rank efficiency is better characterized as a resource profile than by a single metric. Under matched FLOPs, low-rank weight methods can trade lower per-token cost for more training data and close much of their quality gap to full-rank training. Memory savings depend on which components are reduced, while FLOP reductions translate into wall-clock speed more effectively as model scale grows. Optimizer-state methods show a different profile, with modest memory savings and additional execution overhead despite nearly unchanged FLOPs. During serving, direct execution and V/O absorption favor different workloads. V/O absorption reuses the trained low-rank factors to reduce decoding latency and memory at large batch sizes without retraining. LR-EPB provides resource-specific guidance for selecting low-rank pretraining methods under different training and serving constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.