acceptodds
Under review as a conference paper at ICLR 2027

Loce: Locality-Aware Lossless Compression for Efficient LLM Inference

Abstract

Lossless compression offers a promising direction for alleviating the memory and bandwidth bottlenecks in LLMs due to inherent floating-point exponent redundancy. We observe that this redundancy primarily arises from extreme value locality and temporal locality, yet existing methods fail to unify these two forms of locality, leaving significant compression potential untapped. To bridge this gap, we propose Loce, a locality-aware lossless compression framework that leverages an algorithm-system co-design. At the algorithmic level, we present a fixed-length Rice-based lossless coding scheme that transforms both value and temporal locality into compact residual representations for lightweight encoding and decoding. At the system level, we introduce a GPU-efficient design with optimized memory layout, decoding kernel, and inference pipeline to accelerate compressed LLM inference. Extensive experiments on various LLMs and floating-point precisions show that, beyond achieving competitive compression ratios, Loce delivers an average 2.3× kernel-level acceleration and 1.5× end-to-end speedup over previous lossless compression methods. Our code is available on Anonymous GitHub.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.