Allocate, Don't Just Rank: Reducing Redundancy in Test-Time Search for Scientific Discovery
Abstract
Test-time search for scientific discovery reuses previously generated programs and mathematical constructions as context for further language-model generation. Yet selecting solutions solely by their individual promise can concentrate a fixed search budget on structurally similar solutions, producing redundant exploration. We formulate solution reuse as batch allocation: deciding which solutions should be explored together to advance the discovery frontier. We introduce LC-DFT, a submodular optimization framework combining distributional frontier-tail quality (DFT) with lineage coverage (LC). DFT scores the best outcomes of previous rollout bundles against historical rewards, providing a dense quality signal even when frontier improvements are rare. LC uses shared search-tree ancestry to discount redundant contributions within the selected batch. The resulting search policy applies with either a fixed or a continually updated language model. With Qwen3-8B, LC-DFT improves mean performance on mathematical construction, scheduling, and denoising tasks relative to the PUCT-based baseline, while DFT achieves the lowest mean GPU-kernel latencies. Search-tree analysis shows reduced within-batch ancestry overlap alongside later breakthroughs from solutions reused across steps. With Qwen3.8-27B, LC-DFT establishes a new state of the art in Circle-32 packing and produces high-performance GPU kernels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.