ADEPT: Action-Sensitive World Action Models with Adaptive Deliberation for Robotic Manipulation
Abstract
World action models (WAMs) compare candidate actions through imagined futures, but effective manipulation requires predictions that distinguish action consequences and computation allocated to decisions that benefit from deliberation. Generic visual prediction can emphasize nuisance variation, recursive errors can obscure candidate differences, and progress-only scoring can favor inaccurate futures. We present ADEPT, an action-sensitive latent WAM that learns what to predict for action selection and when to plan. A learned token mask preserves action-dependent information; multi-step autoregressive training aligns latent prediction with the deployment horizon. A progress-informed gate invokes candidate evaluation, which combines predicted task progress with rollout consistency. Clean-only component controls and same-state candidate comparisons test representation, recursive training, and action ranking separately. ADEPT achieves 99.1% success on LIBERO and 88.1% Macro-7 on LIBERO-Plus after standard-LIBERO-only training, exceeding its frozen action prior by 16.6 percentage points. On RoboCasa365 it achieves 25.5% split-macro success. Across six real-robot tasks, ADEPT improves pooled success from 55.3% for matched frozen-prior controls to 90.3%. Adaptive deliberation reduces average decision latency by 31% relative to always-on planning in the profiled LIBERO-Spatial setting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.