acceptodds
Under review as a conference paper at ICLR 2027

RarePlay: Exploring Scarce Action Spaces in Tool Graphs via Self-Evolving Synthesis

Abstract

Large language models (LLMs) have demonstrated strong function-calling capabilities, yet their performance on long-horizon tool-use tasks remains heavily dependent on the quality and diversity of training trajectories. Existing synthetic data generation methods typically rely on static tool graphs or multi-agent collaboration, which can produce homogeneous trajectories, incur substantial computational overhead, and may struggle to adapt task difficulty to the evolving capabilities of the model. We introduce RarePlay, a self-evolving framework for synthesizing diverse and progressively challenging function-calling data. RarePlay constructs a state transition graph from the current model’s behavior and introduces Rarity-Guided Exploration (RGE) to prioritize actions that are underrepresented under the current policy while remaining likely to lead to successful outcomes. Sampled trajectories are converted into training tasks through inverse task synthesis and further filtered using rubric-based rejection sampling to improve logical consistency and trajectory quality. By iteratively updating both the model and the state graph, RarePlay incorporates newly explored states and transitions, providing a richer foundation for sampling longer paths and supporting an easy-to-hard curriculum guided by a predefined hop-count schedule. Experiments on BFCL-v4 and AutomationBench demonstrate that RarePlay improves function-calling performance and achieves state-of-the-art results under comparable training-data budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.