Scoring Signal Selection for Test-Time Scaling
Abstract
Test-time scaling (TTS) improves reasoning by sampling multiple candidate solutions and selecting among them with a scoring function. Scoring signals include internal signals (e.g., model confidence), which add negligible cost beyond generation, and external signals (e.g., process reward models, PRMs), which provide additional evaluation at extra cost. Prior work largely compares or combines such signals under a fixed scoring rule across instances. We ask whether the scoring signal can be selected per instance after candidate generation but before invoking external scoring. To our knowledge, this post-generation selection problem has not been explicitly studied. In this work, we systematically study post-generation scoring signal selection. We empirically show that internal and external signals exhibit low rank correlation when evaluating the same candidate sets, while each uniquely succeeds on a non-trivial subset of instances. This complementarity creates a meaningful opportunity for per-instance signal selection. We then evaluate practical selection strategies, ranging from simple candidate disagreement statistics to learned routing models. We find that inexpensive disagreement statistics such as answer entropy capture useful coarse routing signals, while the learned selectors we evaluate do not consistently improve upon them across the full accuracy–call-rate frontier, even with labeled supervision and richer representations. These findings establish scoring signal selection as a practically relevant and largely unexplored axis for improving TTS efficiency, while showing that fine-grained per-instance selection remains challenging beyond coarse disagreement signals.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.