Bicameral: Memory-Efficient Transformer via Latent Tensor Rematerialization
Abstract
Memory capacity limits the size of language models that can be trained and served on given hardware. Quantization lowers the precision of each weight, but leaves the number of trainable parameters unchanged, and with it the gradients and optimizer states that dominate training memory. We introduce Latent Tensor Rematerialization (LTR), which generates each weight matrix on demand from compressed latent codes with small learned generators, and in doing so converts a standard matrix into a recomputable activation. This reduces memory usage by directly storing fewer parameters and optimizer states, as an alternative to lowering precision. Moreover, LTR leverages underutilized compute to reduce the number of parameters that must be stored, and its full-precision parameters remain compatible with quantization. We validate the effectiveness of LTR with Bicameral, a LLaMA-style transformer architecture whose weight matrices are generated on demand by small decoders. %When trained to 10B tokens, a Bicameral model with half the stored parameters of a 2B-parameter transformer matches its validation perplexity and outperforms a standard transformer with the same number of stored parameters. We train Bicameral models with 2B effective parameters to 10B tokens, and find that Bicameral matches the uncompressed baseline on perplexity and zero-shot benchmarks, while reducing stored parameters by 50% and training memory usage by up to 48%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.