PICO: Lossless Compression Stacked on Pruning and Quantization for Efficient LLM Inference
Abstract
Large Language Models (LLMs) are memory-bound during inference, where memory access dominates latency and energy consumption, making the combination of pruning and quantization essential for efficient deployment. Unstructured pruning stays near dense accuracy up to about 50% sparsity without retraining, whereas structured pruning degrades sharply at such sparsity without fine-tuning, but storing the irregular positions of the remaining weights limits its memory savings. Previous XOR-based compression regularizes such sparse weights into fixed-size blocks, yet at this sparsity it generates excessive correction patches that offset its gains, and shift-register-based variants reduce patches at the cost of substantial hardware and encoding overhead. We propose PICO (Patch-minimized Inverted XOR COmpression), a lossless, training-free scheme that stacks on existing pruning and quantization pipelines. PICO combines a hardware- and encoding-efficient XOR map with grouped inverted-XOR encoding, which selects the best map per group and skips patch metadata for mismatch-free groups. PICO applies regardless of unstructured pruning criterion or quantization format while preserving the regular, fixed-size access patterns required for parallel hardware decoding. Across LLaMA2/3.1, Phi-4, and Qwen2.5/3, PICO delivers 11.9% additional memory reduction at 50% sparsity with the smallest XOR decoder, 31% less area and 30% less power than the strongest prior scheme, while encoding 5.9× faster, and its sequential variant PICO-S doubles the reduction of that scheme.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.