FEED: Feedback-Guided Self-Training and Preference Optimization for Efficient Test-Time Discovery
Abstract
Large language models (LLMs) increasingly support scientific discovery through test-time learning, but generating candidate solutions is costly, and high-quality examples remain scarce. To improve learning under limited computational budgets, we introduce FEED (Feedback-Enhanced Efficient Discovery), which turns evaluation feedback into a unified objective for quality-weighted self-training and prompt-matched success–failure preference optimization. FEED draws valid programs from current and historical evaluations, assigning each a uniform base weight and additional emphasis according to solution quality. The same weights scale preference supervision against failures generated under the same original prompt, while every selected valid program retains direct supervision, including those without matched failures. Evaluated with Qwen3-8B on six tasks spanning mathematical construction, heuristic design, and function optimization, FEED achieves earlier discovery of high-quality solutions on multiple tasks than strong baselines, including TTT-Discover, under matched one-hour budgets covering generation, evaluation, and learning. For example, on CVRP-ACO, FEED reaches the 10.2 route-cost threshold in all five runs versus none for TTT-Discover, reducing the one-hour restricted mean time to threshold from 60.00 to 17.62 minutes. Further circle-packing evaluations show efficiency gains across Qwen3 model sizes and solution-quality improvements in two other model families. Together, these results demonstrate that FEED's feedback-guided joint optimization of self-training and success–failure preferences enables earlier discovery of high-quality solutions under fixed test-time budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.