acceptodds
Under review as a conference paper at ICLR 2027

Is Test-Time Video Generation Unnecessary? Adaptive Video Reasoning for World Action Models

Abstract

Recent World Action Models (WAMs) remove the future video prediction at test time, which leads to lower inference cost. This raises a question: is test-time video generation unnecessary? We empirically find that removing it can easily degrade control performance, while its benefit varies substantially across decision states and most of the performance gains can be achieved with limited predictive computation. Therefore, we propose ThinkWAM, a world action model with adaptive predictive reasoning, equipped with a router that dynamically determines whether and how much future prediction to perform for each decision state. To support adaptive predictive reasoning, we formulate a multi-intensity reasoning training scheme that enables one WAM to support both direct action prediction and predictive reasoning at different intensities. We then train the router in two stages. First, it learns the benefit of predictive reasoning; second, it learns a cost-aware allocation over reasoning intensities by balancing predicted control quality and computation cost. At test time, the router determines whether predictive reasoning is needed and, when necessary, selects an appropriate reasoning intensity. ThinkWAM achieves success rates of 98.8%, 80.5%, and 77.7% on LIBERO, LIBERO-Plus, and RoboCasa GR1, respectively and 73.8% on real-world tasks. With predictive reasoning applied to only 28% of decision states, ThinkWAM recovers up to 68% of the success-rate gain achieved by reasoning at every state.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.