How Can Every Rollout Count for Skill Seekers? COMPASS for Skill Proposers
Abstract
Textual-program optimization must balance proposing and evaluating candidates under a fixed rollout budget. Evaluating each candidate more extensively leaves fewer rollouts for proposing new candidates. We introduce COMPASS, which turns accumulated evaluation results into relative credit to guide subsequent proposals. On each shared instance, attaining the highest observed score earns more credit when a smaller fraction of the other candidates attain it. COMPASS averages these per-instance credits to select candidates for further revision and updates the credits as new results arrive. For revision, it assigns batches of instances using the selected candidates’ score gaps to the best score recorded among them on each instance. Each new candidate is evaluated on instances that were not used to generate it or its ancestors. Across six benchmarks, COMPASS achieves the highest average score among the compared methods for both Qwen3-8B and GPT-5.4-nano. On Qwen3-8B, it improves over GEPA by 11.8 percentage points on average and 36.7 points on AIME-2025. Ablations on Qwen3-8B show that both relative credit and matching improve the average score across the six benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.