Joint Multi-View Latent Actions for Robust Vision-Language-Action Learning
Abstract
A latent action model (LAM) maps a pair of video frames to a discrete latent action that a vision-language-action (VLA) policy learns to predict, so that the policy can learn from videos without action labels. Existing LAMs infer the latent action, mostly, from one camera, whose frames record the physical change only as that camera projects it, together with camera-specific content; the latent action can therefore lose part of the change and carry that content instead. JMV-LA infers the latent action from all cameras a recording provides: a shared transition encoder summarizes every available camera's frame pair, learned queries read all summaries before quantization, and one shared decoder must predict every camera's next features from the single latent action and that camera's first frame, so that the latent action has to supply the change in a form that serves every projection. The VLA predicts it from the primary camera alone and needs no additional camera at deployment. Compared with LAMs learned from the primary camera only, JMV-LA yields lower action-decoding error and a readout that is more stable under image perturbations. The single-camera VLA supervised by JMV-LA achieves the highest average success among the compared policies on the four LIBERO suites, LIBERO-10 camera-pose perturbations, and LIBERO-Pro, as well as the highest observed success in six of seven ALOHA real-robot evaluation settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.