acceptodds
Under review as a conference paper at ICLR 2027

TempQuant: Learning Adaptive Temperatures For Quantized Language Models

Abstract

Large language models are increasingly deployed at low precision for applications ranging from reasoning to creative writing and beyond, yet quantization can degrade their performance. These applications often use different decoding settings, including sampling temperature and top-p, which control the randomness of token selection. We study the interaction between quantization and these settings by examining how quantization distorts output token probabilities. These distortions can change generated responses and affect task performance. We characterize the connection between quantization and temperature by measuring changes in generation-time entropy across deployed temperatures, and examine whether these changes can be understood as a temperature-like shift. We also investigate how much distributional error can be corrected through our intervention method, TempQuant. TempQuant is a compact learned predictor that adaptively sets the quantized model’s temperature and top-p value at each decoding step to minimize divergence between the full-precision and quantized probability distributions. At deployment, it uses only features of the quantized model’s out- puts, requiring no full-precision reference. We show that TempQuant can improve model accuracy and analyze how these corrections affect model behavior during generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.