acceptodds
Under review as a conference paper at ICLR 2027

TAM-WM: Transition-Aligned Multi-View World Model for Visual Robot Manipulation

Abstract

Multi-view observations provide complementary information for visual robot manipulation, but they also require a world model to represent the consequences of an action consistently across viewpoints. Existing approaches often encourage synchronized observations to have similar representations, which primarily enforces agreement about the current state. For action selection, however, the model must also assign consistent finite-horizon changes to the same physical transition; strong state agreement alone does not guarantee such transition agreement. To address the issue of inconsistent transition representations across viewpoints, we propose TAM-WM, which learns a shared visual predictive coordinate in two stages. Stage A aligns the displacements induced by the same observed physical transition across synchronized views. Stage B freezes this coordinate and learns state- and action-conditioned residual dynamics, enabling candidate actions from a shared current state to be compared through their predicted changes. A downstream online correction teacher uses these predictions to score local proposals and selectively correct a nominal policy action. On Can PH, a fixed-scorer ablation shows that removing transition alignment reduces demonstrated-action ranking, while the learned scores correlate with simulator-executed candidate outcomes. Across eight simulated manipulation tasks, the complete system achieves the highest equal-weight average success-rate point estimate among the evaluated methods: 75.85%, 5.28 percentage points above VILA and 8.99 points above the behavior-cloning anchor.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.