acceptodds
Under review as a conference paper at ICLR 2027

InterACT: Grounding Latent Actions through Camera–Physics Interventions

Abstract

Latent action models learn transitions from video, but visual changes reflect both physical interaction and camera motion. Temporal consistency does not guarantee camera invariance. We introduce , which adapts a pretrained latent action model through crossed camera and physical interventions. Each quartet pairs two interactions with two camera trajectories. A bounded residual aligns views, a frozen reference anchors physical distance scale, and separate gradient paths train the correction and encoder. achieves lower camera ratios and action errors and higher outcome gains than video continuation and an existing ego-supervised adaptation across three controlled domains. On unseen camera trajectories, it lowers the camera-to-physical distance ratio from to while retaining near-reference physical scale, and reduces outcome prediction error by relative to a state-only probe. Controlled cross-view Recall@5 increases from to . A separate comparison with a supervised baseline reveals a trade-off: the full method better preserves reference scale, while supervised adaptation yields higher outcome gains. For downstream policy learning, the frozen tokenizer yields success with on MetaWorld MT50 and with ABot-M0.5 on RoboCasa365 atomic-seen tasks, exceeding the reported references of and , respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.