acceptodds
Under review as a conference paper at ICLR 2027

AnyTool-Agent: API-Guided Tool Use for Data-Scarce Ecosystems via Execution-Grounded Policy Optimization

Abstract

Tool-using agents can perform well with familiar APIs, but adapting them to a new API ecosystem presents two challenges. First, demonstrations involving unfamiliar APIs are scarce, making it difficult to obtain successful trajectories for an initial policy. With binary outcome rewards, GRPO assigns zero advantages when all rollouts in a group fail, providing no outcome-based learning signal. Second, Vanilla GRPO applies a shared trajectory-level advantage across generated tokens, failing to distinguish necessary, redundant, and erroneous API calls within a multi-step trajectory. API specifications and existing environment entities guide demonstration construction, while sandbox feedback supports action-level credit assignment. We introduce AnyTool-Agent, which combines API-Guided Data Synthesis with Execution-Grounded Policy Optimization (EGPO). API-Guided Data Synthesis selects API combinations using their specifications and grounds them in existing environment entities to construct tasks bottom-up. Expert trajectories passing semantic completion checks provide supervised fine-tuning demonstrations, while their core API sequences, successfully re-executed in fresh sandboxes, provide verified references for reinforcement learning. EGPO scores actions through reference matching and deterministic execution feedback, with within-group scarcity weighting for infrequently visited reference APIs. These scores multiplicatively reweight the outcome advantage to produce action-level advantages, preserving its sign under positive reweighting factors while distinguishing action contributions. Under the same task-generation budget, API-Guided Data Synthesis covers as many distinct APIs as random exploration. On AppWorld and BFCL, EGPO improves avg@4 over Vanilla GRPO by 4.8–11.2 percentage points across two Qwen3 backbones with identical SFT initialization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.