Behavior-Aligned Neural Supervision for Vision-Language-Action Models
Abstract
Vision-language-action (VLA) models learn embodied representations primarily from externally observable signals such as vision, language, and action trajectories. However, external observations may not fully capture behaviorally relevant structure and temporal dynamics during physical interaction. Biological organisms and embodied agents both operate through continuous perception–action loops under related physical constraints, motivating us to ask whether neural activity recorded during human interaction can provide complementary supervision for embodied learning. To answer this, we investigate EEG as a training-time supervisory signal for adapting pretrained embodied models. We first synchronously collect egocentric video and 128-channel EEG during natural manipulation tasks. We then introduce EEG-guided representation learning, with EEG used only during training and removed at inference. Using OpenVLA as the primary backbone, we evaluate the approach across multiple manipulation benchmarks and out-of-distribution settings involving visual, semantic, and execution shifts, and further validate its transfer to on LIBERO. EEG-guided training improves task success over the corresponding robot-only baselines across these settings and backbones. These gains depend on aligned EEG–behavior correspondence, as they are reduced when the alignment is disrupted or EEG supervision is ablated. Behavioral and representation analyses further associate the improvements with more effective task progression and completion, with the strongest representational changes occurring in late, action-proximal layers. These results suggest that behavior-aligned human neural dynamics can provide transferable training-time supervision for embodied representation learning without introducing an additional modality at deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.