RarePlay: Exploring Scarce Action Spaces in Tool Graphs via Self-Evolving Synthesis
Abstract
Large language models (LLMs) have demonstrated strong function-calling capabilities, yet their performance on long-horizon tool-use tasks remains heavily dependent on the quality and diversity of training trajectories. Existing synthetic data generation methods typically rely on static tool graphs or multi-agent collaboration, which can produce homogeneous trajectories, incur substantial computational overhead, and may struggle to adapt task difficulty to the evolving capabilities of the model. We introduce RarePlay, a self-evolving framework for synthesizing diverse and progressively challenging function-calling data. RarePlay constructs a state transition graph from the current model’s behavior and introduces Rarity-Guided Exploration (RGE) to prioritize actions that are underrepresented under the current policy while remaining likely to lead to successful outcomes. Sampled trajectories are converted into training tasks through inverse task synthesis and further filtered using rubric-based rejection sampling to improve logical consistency and trajectory quality. By iteratively updating both the model and the state graph, RarePlay incorporates newly explored states and transitions, providing a richer foundation for sampling longer paths and supporting an easy-to-hard curriculum guided by a predefined hop-count schedule. Experiments on BFCL-v4 and AutomationBench demonstrate that RarePlay improves function-calling performance and achieves state-of-the-art results under comparable training-data budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.