acceptodds
Under review as a conference paper at ICLR 2027

Screen-Then-Select: Learning When Temperature Selection Helps LLM Reasoning

Abstract

Sampling temperature controls the diversity of language-model reasoning, yet aggregate accuracy can conceal substantial differences in how inputs respond to temperature changes. We characterize this heterogeneity through multi-model MATH sweeps and high-sampling continuation experiments, connecting modest aggregate changes in a conventional temperature range to large local opportunities. Motivated by these findings, we propose Screen-Then-Select, an answer-assisted framework that separates screening for responsive candidates from selecting their temperatures. A small pool of reference-checked responses identifies mixed-success questions; a lightweight prompt-only selector then chooses a temperature for each fresh answer without updating the language model. In a MATH prefix study, correct-answer counts enrich inputs with large temperature responses. Trained on MATH and frozen for transfer, the selector improves individual-answer accuracy on screened GSM8K and OlympiadBench subsets relative to fixed temperatures selected for that metric on separate development questions, including a 7.1-percentage-point gain on GSM8K. Matched rejected controls show smaller observed gains, while policy decomposition identifies a shared temperature change as the source of the GSM8K improvement. Together, these results provide an empirical foundation and a lightweight framework for targeting temperature selection where its observed benefits are concentrated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.