Beyond Sampling More: Scaling LLM Agents through Skill-Conditioned Exploration
Abstract
Repeated sampling is a simple and effective form of test-time scaling for LLM agents, yet all candidate trajectories are typically generated under the same procedural context. We investigate reusable skills—packaged procedural guidance—as a source of structured exploration and study how a fixed rollout budget should be allocated between breadth across skills and depth within a single skill. Using a provenance-tracked pool of 4,802 open-source skills, we introduce Skill-Diverse Best-of-N (SD-BoN), a label-free candidate generation method that selects eight skills via a relevance–diversity criterion and assigns sixteen rollouts to each. Across four models evaluated on AppWorld, WebArena, GAIA, and tau-bench, SD-BoN with flat majority voting consistently outperforms similarity-based Top-1 skill selection, improving performance by 8.33 points for GPT-5.6-sol on AppWorld and by 6.96–7.51 points on the remaining benchmarks, with a mean gain of 7.26 points across all 16 model–benchmark pairs (range 5.22–9.57). In contrast, a label-aware pilot-ranking diagnostic that uses 1,024 success-labeled screening rollouts per task and model produces a smaller mean allocation gain of 1.85 points. Our analysis further distinguishes candidate coverage from selected success: broader skill exploration reliably increases the availability of successful candidate trajectories, but also enlarges the gap that a practical selector must bridge. These results show that evaluating both coverage and selection effectiveness, rather than coverage alone, is essential for understanding skill-conditioned test-time scaling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.