acceptodds
Under review as a conference paper at ICLR 2027

IA-World: Residual VLA Post-Training with Latent World Model

Abstract

Vision-language-action (VLA) models provide strong priors for robotic manipulation, yet adapting them efficiently through online reinforcement learning remains challenging. Existing post-training methods often correct base actions using current observations or policy-derived signals. However, they typically do not explicitly anticipate the consequences of the sampled action plan. Meanwhile, sparse task rewards provide limited feedback on whether a correction actually improves future task progress. In this work, we revisit VLA post-training from the perspective of action-specific foresight and argue that effective residual adaptation should anticipate the consequences of a base plan before modifying it, while learning from how that anticipated future changes after execution. Based on this insight, we propose IA-World, a framework that uses a latent world model for residual VLA post-training. Given an observation, a frozen VLA first proposes a base action plan; we distill its action representations associated with the plan into an Intent-Action Token (IA-Token) and use a lightweight latent world model to predict the multi-step consequences of executing the uncorrected plan. The predicted future directly conditions a residual policy to correct the base action before execution, forming a control loop of observing, proposing, predicting, correcting, and executing. After execution, a foresight scorer conditioned on the goal reevaluates the predicted future from the new state, and the resulting change in task progress provides dense learning signals beyond sparse environment rewards. We evaluate IA-World across LIBERO and ManiSkill with both Isaac GR00T and OpenPI. IA-World consistently improves task success through residual correction and outperforms competing post-training methods based on reinforcement learning. The results demonstrate that modeling the future consequences of a VLA’s own action plan enables a shift from reactive residual correction to post-training guided by foresight.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.