Sampling at Infinite Temperature: KL* and KL-Budget Decoding
Abstract
Applying temperature before truncation lets unlikely tokens into the candidate set, and a high temperature can then derail generation. We reverse the order. The model’s original probabilities choose the candidates, and exploration happens only among them. At infinite temperature every candidate is equally likely, so the candidate set alone decides what is sampled. Global KL* picks the candidate set whose uniform draw is closest to the model in KL divergence. We also study local KL*, a greedy variant, and KL-budget, which flattens a min-p candidate set within a fixed KL allowance. On the complete GSM8K test set, both KL* rules match each model’s recommended sampling settings without any tuning, and global KL* has no significant loss in pass@8 or majority vote against any standard default sampler. Across six tasks, the KL* rules gain two to three times as often as they lose against these defaults. On real next-token distributions they explore below the top token while rarely reaching the tail, and at high temperature truncating first prevents the collapse of high-temperature sampling. The rules remain competitive in creative writing and with reasoning models. We release all generations and analysis code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.