acceptodds
Under review as a conference paper at ICLR 2027

Composition Discovery in Tool Worlds by Active Exploration

Abstract

Agents usually discover connections between tools, which we call handoffs, while solving a task, leaving the knowledge tied to specific trajectories. We investigate whether an agent can instead explore a tool world, a set of callable functions together with the data behind them, before any task, with the explicit goal of building reusable knowledge about its tools. Task-free exploration faces two bottle-necks: the combinatorial number of candidate handoffs, and the volatility of tool data. Because experiments require valid inputs from earlier calls, every change of the data invalidates the inputs collected so far. We introduce ICHNAEA, an explorer that maintains a Beta belief over candidate handoffs and selects the calls whose outcomes are expected to be most informative. To overcome data volatility, ICHNAEA stores each successful call as a recipe and, when the data change, re-executes its recipes to rebuild fresh inputs. We evaluate ICHNAEA on procedural tool worlds, where the true handoffs are known and names can be hidden, and on AppWorld, a benchmark of simulated everyday apps for language-model agents. In procedural worlds, ICHNAEA learns handoffs that complete every target workflow in unseen data, with no false handoff, while without rebuilding its inputs the explorer completes 0–73% depending on how much the names reveal. A language-model agent given these handoffs reaches every target, against 0–89% without them. On AppWorld, documentation alone already lets the agent reach nearly every goal, and some calls succeed only by coincidence. This points to two limits outside generated worlds: what exploration adds depends on how much of a suite’s composition its documentation leaves unstated, and without known true handoffs, discovery is hard to measure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.