Is Test-Time Video Generation Unnecessary? Adaptive Video Reasoning for World Action Models
Abstract
Recent World Action Models (WAMs) remove the future video prediction at test time, which leads to lower inference cost. This raises a question: is test-time video generation unnecessary? We empirically find that removing it can easily degrade control performance, while its benefit varies substantially across decision states and most of the performance gains can be achieved with limited predictive computation. Therefore, we propose ThinkWAM, a world action model with adaptive predictive reasoning, equipped with a router that dynamically determines whether and how much future prediction to perform for each decision state. To support adaptive predictive reasoning, we formulate a multi-intensity reasoning training scheme that enables one WAM to support both direct action prediction and predictive reasoning at different intensities. We then train the router in two stages. First, it learns the benefit of predictive reasoning; second, it learns a cost-aware allocation over reasoning intensities by balancing predicted control quality and computation cost. At test time, the router determines whether predictive reasoning is needed and, when necessary, selects an appropriate reasoning intensity. ThinkWAM achieves success rates of 98.8%, 80.5%, and 77.7% on LIBERO, LIBERO-Plus, and RoboCasa GR1, respectively and 73.8% on real-world tasks. With predictive reasoning applied to only 28% of decision states, ThinkWAM recovers up to 68% of the success-rate gain achieved by reasoning at every state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.