ParallelWorld: Test-Time Scaling for Agentic Embodied Reasoning
Abstract
Embodied AI has emerged as a key paradigm for autonomous systems operating in complex physical environments. While end-to-end models have established strong local capabilities for perception and control, scaling these capabilities to extended, multi-stage tasks remains difficult. Agentic embodied systems address this challenge by introducing high-level planners that organize perception, reasoning, and executable skills over longer interaction horizons. However, existing agentic embodied frameworks typically generate actions for direct execution without explicitly simulating how alternative action sequences shape future observations and task progress, leading agents to commit to suboptimal interactions before considering their long-term consequences. To address this limitation, we propose ParallelWorld, an adaptive-horizon test-time scaling framework that grounds agentic embodied reasoning in the prospective simulation of action outcomes. Specifically, we design an agent system to simulate and evaluate alternative trajectories in parallel before committing to an action. A verifier assesses intermediate state transitions with respect to the task objective, dynamically pruning unpromising branches and guiding subsequent expansion. Based on the evaluated long-horizon outcomes, the agent selects a trajectory, executes only its first action, and replans as new observations become available. Experiments on ESI-Bench and LIBERO-Pro demonstrate improvements in active spatial reasoning and language-conditioned manipulation, respectively, highlighting the effectiveness of prospective planning for both understanding and manipulation tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.