LCT: Losslessly Compressed Tensors with GPU Friendly Execution for Training
Abstract
Lossless tensor compression could reduce memory usage during training, but training requires tensors to be repeatedly re-encoded. We introduce Losslessly Compressed Tensors (LCT), a GPU execution framework designed for this repeated encode–decode regime. LCT exploits the natural redundancy in floating point data to reduce storage for weights, activations and optimizer states. Highly optimized GPU kernels implement Huffman coding with precomputed codebooks for low compression and decompression overhead, essential for training. Our analysis characterizes LCT's storage costs and adaptation to magnitude shifts. LCT achieves a compression ratio of up to , and is demonstrated in training and finetuning settings. On a 162M-parameter NanoGPT, LCT reduces memory usage by for a step-time increase. In Qwen3-4B fine-tuning, LCT saves up to of peak memory, enabling a 3.5X larger context length in training on the same GPU.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.