DIWA: Decision-Influential World Abstraction for VLA-WAM Policies
Abstract
Predicting future observations can improve robot decision making, but a controller need not model every visible change with equal fidelity. We introduce Decision-Influential World Abstraction (DIWA), a framework that allocates future modeling capacity according to its effect on action generation. A lightweight estimator ranks object–time queries before an action-conditioned world decoder expands them. Its supervision measures changes in action responses, value estimates, and task progress under latent interventions, with shared randomness for stochastic action heads. A complementary regret geometry organizes representations by the relative returns of aligned candidate action rules. During inference, only selected future queries are decoded, avoiding dense future generation. Across LIBERO, RoboTwin, RoboCasa, and real-robot manipulation, DIWA achieves 73.5% aggregate success, improving over DreamVLA by 5.6 percentage points. The out-of-distribution evaluation yields 75.8% success, a 10.6-point improvement. Mean inference latency is 89 ms, compared with 213 ms for dense imagination and 241 ms for WorldVLA. Three-seed simulation results, 300 physical trials per method, budget sweeps, and intervention diagnostics distinguish the roles of influence supervision, regret geometry, and selective computation. These results support using decision sensitivity to determine which future information a robot policy should compute.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.