ACE-Sampler: Hierarchical Stochastic Data Selection for Tool-Using Agents
Abstract
Training tool-using language models increasingly relies on heterogeneous collections of agent trajectories, yet larger data pools can introduce redundancy, imbalance, and unreliable supervision. Selecting useful agentic data requires capturing state-dependent behaviors and accounting for the uneven learning value of turns within a trajectory. We introduce ACE-Sampler, a hierarchical stochastic sampler that jointly considers Accuracy, Complexity, and divErsity under a fixed turn-level training budget. At its core is a factorized behavioral representation that captures interaction patterns and action dependencies beyond surface semantics. Using this representation, ACE-Sampler first samples trajectories for behavioral coverage, then prioritizes important assistant turns and samples them for novelty. A stochastic continuation rule balances coverage across trajectories with supervision within each trajectory, while validity checks and learner-aware complexity estimation favor valid examples of appropriate complexity. Across three tool-use benchmarks and three model scales, ACE-Sampler consistently outperforms the compared selectors with a budget of 5,000 training turns, improving aggregate accuracy over random selection by more than 4 percentage points on average. Ablations and scaling experiments examine the effects of behavioral representations, validity filtering, learner-aware complexity, and training data budgets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.