Token Caps, Not Probes: Managing Non-Termination in Distilled Reasoning Models
Abstract
We study how thinking-token budgets relate to accuracy in DeepSeek-R1-Distill-Qwen-7B on GSM8K, MATH-500, and AIME\@. On GSM8K, accuracy peaks at a very small forced thinking budget, while on MATH-500 short budgets remain well below uncapped accuracy. On AIME, generations divide by whether the model ends its reasoning within a 10,000-token ceiling. Generations that terminate are usually correct; nearly half never terminate and consume most of the inference compute. Whether a problem terminates changes for about a quarter of problems under temperature sampling or across hardware. Forcing termination at a fixed cap requires no predictor and traces a compute–accuracy frontier: at 4,000 tokens it saves about 39% of inference compute with accuracy no lower than generating to the 10,000-token ceiling. Linear probes on early hidden states do not predict non-termination at the primary evaluation point; one position in a nine-position sweep (token 125) is at the threshold for significance after correction for multiple testing; replication is needed. We also show that the correction, valid for unsupervised probes, inflates the performance of supervised ones and produced a spurious positive result in our initial analysis. These results suggest that token-count caps, rather than activation probes, are the practical lever for reducing inference cost in this setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.