DLR: Zero-Inference-Cost Latent Residuals for Low-Rank Pre-Training
Abstract
Large language models have driven recent progress in language and multimodal AI, yet pre-training them at scale is prohibitively expensive. Low-rank pre-training, which factorizes each weight matrix into a rank-r product to reduce both parameters and FLOPs, is a promising response but typically lags behind full-rank training in quality. We propose Duplicated Latent Residual (DLR), a training-only, parameter-free, foldable plug-in for low-rank pre-training. DLR augments the standard low-rank output with a fixed structured residual that expands the latent representation by replicating each latent coordinate across the output dimension. The replication factor is determined by the ratio between the output dimension and the latent rank. With a fixed residual scale, DLR adds zero learnable parameters per layer. After training, the residual is absorbed into the up-projection in closed form, so the deployment parameter count, FLOPs, and memory match the underlying low-rank backbone exactly. Across LLaMA models from 60M to 7B parameters, DLR consistently strengthens low-rank pre-training on C4 validation perplexity; folded checkpoints transfer cleanly to supervised fine-tuning on standard benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.