acceptodds
Under review as a conference paper at ICLR 2027

PICO: Lossless Compression Stacked on Pruning and Quantization for Efficient LLM Inference

Abstract

Large Language Models (LLMs) are memory-bound during inference, where memory access dominates latency and energy consumption, making the combination of pruning and quantization essential for efficient deployment. Unstructured pruning stays near dense accuracy up to about 50% sparsity without retraining, whereas structured pruning degrades sharply at such sparsity without fine-tuning, but storing the irregular positions of the remaining weights limits its memory savings. Previous XOR-based compression regularizes such sparse weights into fixed-size blocks, yet at this sparsity it generates excessive correction patches that offset its gains, and shift-register-based variants reduce patches at the cost of substantial hardware and encoding overhead. We propose PICO (Patch-minimized Inverted XOR COmpression), a lossless, training-free scheme that stacks on existing pruning and quantization pipelines. PICO combines a hardware- and encoding-efficient XOR map with grouped inverted-XOR encoding, which selects the best map per group and skips patch metadata for mismatch-free groups. PICO applies regardless of unstructured pruning criterion or quantization format while preserving the regular, fixed-size access patterns required for parallel hardware decoding. Across LLaMA2/3.1, Phi-4, and Qwen2.5/3, PICO delivers 11.9% additional memory reduction at 50% sparsity with the smallest XOR decoder, 31% less area and 30% less power than the strongest prior scheme, while encoding 5.9× faster, and its sequential variant PICO-S doubles the reduction of that scheme.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.