acceptodds
Under review as a conference paper at ICLR 2027

CARVE: Learning Latent Actions under Cross-View Visual Heterogeneity

Abstract

Latent action models learn action-centric dynamics from unlabeled robot videos, but existing methods predominantly operate on a single camera stream. Recent multi-view formulations instead emphasize view invariance, overlooking that synchronized cameras provide complementary but not interchangeable evidence about the same physical action. We introduce CARVE (Cross-view Action Representation with View-specific Evidence), a latent action framework that selectively shares transition information across synchronized views. For each target view, CARVE forms global components from synchronized multi-view context and retains target-view private components. A shared future-dynamics model routes global information across target views while restricting private information to its corresponding predictor. Because synchronized multi-view videos are substantially less abundant than single-view recordings, CARVE uses a progressive three-stage procedure that combines single-view pretraining, amplitude-guided multi-view adaptation, and global–private decomposition. CARVE consistently improves long-horizon video prediction and downstream robot control over strong latent action baselines. Representation analyses and ablations further show that selective cross-view sharing captures complementary action information while preserving target-specific transition cues.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.