acceptodds
Under review as a conference paper at ICLR 2027

One View, One Action: Spatially Generalizable VLAs through Observation-Aligned Experts and Cross-Graph Consistency

Abstract

Multi-view vision-language-action (VLA) models typically learn actions only in the robot's control frame, requiring prediction to account for the coordinate mapping between each camera and the robot. We propose **One View, One Action (OVOA-VLA)**: each view learns an action target expressed in its own camera coordinates. The same physical movement thus provides distinct, observation-aligned supervision for each view. We pair these targets with robot states in the corresponding frames to learn the motion required by the current robot–scene configuration, targeting spatial generalization across both camera viewpoints and robot states. OVOA-VLA jointly trains one flow-based action expert per view and a base-frame execution expert. Two complementary directed attention graphs allow the per-view experts' action losses to train the base representation and allow the base expert to use their action features. Cross-Graph Consistency encourages agreement between the two graphs' base-frame predictions for the same sample. The model supports inference with all experts or the base expert alone. Experiments on LIBERO, LIBERO-Plus, VLABench, and MimicGen show improved performance. On LIBERO-Plus, OVOA-VLA achieves 88.8% success, with gains of 9.2% and 21.2% over under camera and robot-initialization perturbations, respectively. Base-only inference retains 86.5% overall success, showing that observation-aligned supervision benefits the executable policy after removing the per-view experts. On real-world, long-horizon, out-of-distribution tasks, it outperforms by 8.3%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.