WorldForge: Benchmarking Persistent Embodied Agents in Living Worlds
Abstract
Embodied agents must often act with incomplete world state, yet missing information can be handled in fundamentally different ways: an agent can spend more computation reasoning over what it already knows, or act to acquire new evidence. We introduce WorldForge, a simulated household platform for studying such decisions under controlled, replayable world conditions. WorldForge provides generated homes, certified physical interactions, matched task instances, and a shared interface for embodied agents, while supporting changes in information, task structure, and world dynamics. As one controlled study within this platform, we examine everyday object relocation and vary only whether target locations are given or must be recovered through interaction. Across two model–interface systems and multiple reasoning effort settings, withholding locations reduces task completion and induces substantially more search. Increasing fixed reasoning does not reliably close this gap, whereas self-selected reasoning matches the best fixed-effort GPT performance with substantially less computation. Trajectory analysis reveals that most remaining failures occur before the target ever becomes observable, particularly for objects hidden inside closed containers. These results expose a gap between allocating internal computation and allocating information-gathering actions, and illustrate how WorldForge can be used to isolate and study such failures in embodied agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.