HarnessWAM: Bridging Prediction and Deliberation in World Action Models
Abstract
World Action Models (WAMs) jointly learn environmental dynamics and robot actions, introducing priors over physical evolution into embodied control. However, finite-horizon prediction and action generation are insufficient for complex embodied tasks that require global planning, cross-stage state maintenance, execution verification, and failure recovery. We refer to this mismatch as the prediction-deliberation gap. To address this gap, we propose HarnessWAM, an agentic framework for WAMs. HarnessWAM employs a vision-language-model-based Task Manager to maintain an evidence-grounded scene belief and organize the global instruction into a structured task graph. Capability-conditioned executable-space projection then compiles the graph into WAM-supported skill sequences that satisfy task dependencies and embodiment-state constraints. During execution, HarnessWAM couples a lightweight progress estimator with event-triggered Task Manager deliberation. The estimator supplies frequent execution cues, while the Task Manager uses observations, task state, and interaction history to verify outcomes and decide whether to advance, acquire evidence, revise the plan, or recover. Local recovery preserves acquired scene knowledge so that execution can resume after failure. HarnessWAM achieves state-of-the-art average performance on RoboMemArena, with full-task and subtask success rates of 64.4% and 76.8%, respectively, and reaches 23.7% success on RoboCerebra Ideal. On a physical UR3 arm, full-task success reaches 95% for shape-conditioned tabletop tidying and 90% for memory-dependent box exploration. These results demonstrate the effectiveness of model-external agentic orchestration in bridging the prediction-deliberation gap for complex embodied tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.