acceptodds
Under review as a conference paper at ICLR 2027

Quantization for Reasoning Preservation in Small LLMs

Abstract

Large Language Models (LLMs) have demonstrated exceptional capabilities in textual inference, but their practical deployment is constrained by high computation and memory overhead. To address this, we propose LLM-QRP (Quantization for Reasoning Preservation), a layer-aware, mixed-precision quantization framework designed to reduce model memory footprint while preserving multi-step reasoning capabilities. Unlike standard uniform post-training quantization methods that treat all parameters equally, LLM-QRP identifies and protects critical "thinking" sub-components by analyzing layer-wise Self-Logits Evolution Decoding (SLED) and entropy dynamics across the residual stream. These signals are unified via Principal Component Analysis (PCA) into a data-driven reasoning importance score. We then formulate precision allocation as a 0-1 Multiple-Choice Knapsack Problem solved via Integer Linear Programming (ILP) to assign higher bit-widths to high-sensitivity components under strict VRAM targets. Experiments on sub-500M parameter models across GSM8K, TruthfulQA, and MMLU demonstrate that LLM-QRP achieves an optimal Pareto frontier between model compression and task accuracy, maintaining near-baseline reasoning performance while outperforming uniform 8-bit baselines and alternative quantization methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.