ScopeVLA: Task-Grounded and View-Specialized Spatial Conditioning for Robotic Manipulation
Abstract
Incorporating geometric information into vision-language-action (VLA) models has become an important approach to improving robotic manipulation performance. However, identifying task-relevant spatial evidence and exploiting the distinct roles of third-person and wrist views remain challenging. In this paper, we introduce ScopeVLA, a task-grounded framework for view-specialized spatial conditioning. ScopeVLA employs Global and Wrist queries to extract task-relevant spatial distributions from intermediate vision-language model (VLM) representations. These queries are supervised by view-specific workspace targets derived from end-effector trajectories. These distributions guide two complementary conditioning pathways: the Global prior introduces a confidence-aware attention bias over third-person visual tokens, while the Wrist query assignments spatially modulate multi-level features extracted from a frozen geometry model. The modulated geometric features are projected into additional keys and values for joint attention in the action expert. Both pathways operate at selected action-expert layers while preserving the original layerwise vision-language conditioning. Controlled instruction interventions show that the learned spatial distributions remain stable under task-preserving paraphrases while responding to changes in task instructions. Experiments on two simulation benchmarks and real-world manipulation tasks show that ScopeVLA improves task success over the evaluated VLA baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.