AWAM: Agentic World Action Model for Embodied Intelligence
Abstract
Existing Vision-Language-Action (VLA) models and World Action Models (WAMs) typically rely on end-to-end decision-making architectures without explicit agentic reasoning. Although WAMs provide foresight through future prediction, they lack agentic closed-loop reasoning and remain vulnerable to error accumulation in long-horizon, compositional, and out-of-distribution tasks. To address this limitation, we introduce AWAM, an Agentic World Action Model whose dual-brain architecture comprises an Agentic Reasoning Brain and a World–Action Brain, integrating high-level reasoning, world modeling, and action generation. The Agentic Reasoning Brain reasons over task goals, observations, and retrieved experience to generate and refine structured intents. Conditioned on these intents, the World–Action Brain jointly predicts future world states and executable actions, distilling the predicted physical consequences of these actions into world evidence that is fed back to the Agentic Reasoning Brain. This interaction establishes a unified closed loop in which reasoning and world modeling constrain and refine each other to guide action generation. During training, AWAM generates multiple candidate intents conditioned on the same physical state and evaluates them through simulation rollouts. At inference time, AWAM adopts a single-intent execution strategy and re-deliberates when necessary. Experiments across multiple benchmarks demonstrate AWAM’s superior manipulation performance, with strong capabilities in long-horizon execution, compositional generalization, out-of-distribution robustness, and failure recovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.