acceptodds
Under review as a conference paper at ICLR 2027

ACE-Sampler: Hierarchical Stochastic Data Selection for Tool-Using Agents

Abstract

Training tool-using language models increasingly relies on heterogeneous collections of agent trajectories, yet larger data pools can introduce redundancy, imbalance, and unreliable supervision. Selecting useful agentic data requires capturing state-dependent behaviors and accounting for the uneven learning value of turns within a trajectory. We introduce ACE-Sampler, a hierarchical stochastic sampler that jointly considers Accuracy, Complexity, and divErsity under a fixed turn-level training budget. At its core is a factorized behavioral representation that captures interaction patterns and action dependencies beyond surface semantics. Using this representation, ACE-Sampler first samples trajectories for behavioral coverage, then prioritizes important assistant turns and samples them for novelty. A stochastic continuation rule balances coverage across trajectories with supervision within each trajectory, while validity checks and learner-aware complexity estimation favor valid examples of appropriate complexity. Across three tool-use benchmarks and three model scales, ACE-Sampler consistently outperforms the compared selectors with a budget of 5,000 training turns, improving aggregate accuracy over random selection by more than 4 percentage points on average. Ablations and scaling experiments examine the effects of behavioral representations, validity filtering, learner-aware complexity, and training data budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.