acceptodds
Under review as a conference paper at ICLR 2027

Imagine Before You Call: Factored World Models for Tool-Using Agents

Abstract

Tool-using agents extend the capabilities of large language models by interacting with external APIs and environments. However, reliable tool use remains challenging because interaction histories contain substantial task-irrelevant information, while the value of a tool call often depends on consequences that are unavailable before execution. Existing reward models score tool calls directly from the current context without explicitly modeling how each call changes the environment, leaving the agent's understanding of environment dynamics under-constrained. To tackle these issues, we introduce **FOLD**, a factored world-modeling framework for tool-using agents that predicts the consequences of candidate calls through one-step latent dynamics. FOLD partitions the latent state according to two functional properties: (1) whether state information changes in response to the agent's action; and (2) whether it affects future task reward. This yields four latent blocks with distinct transition and decision roles. Based on this representation, FOLD jointly learns latent transitions, process rewards, and state values. The process reward provides step-level supervision for standard policy optimization, while the predicted successor states support test-time look-ahead before a tool is invoked. We further provide a block-wise identifiability analysis for the learned factorization. Experiments on ALFWorld, -Bench, and DataSearch show that FOLD achieves 97.6% average accuracy on pairwise tool-call preference decisions, raises downstream task success by 3.7 points, and improves test-time tool-selection accuracy to 95.7% on average.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.