acceptodds
Under review as a conference paper at ICLR 2027

What to Teach Before Acting: Revisiting Embodied VQA for VLA Models

Abstract

Vision-Language-Action (VLA) models build their policies on pretrained Vision-Language Models (VLMs), and embodied visual question answering (VQA) is widely used to prepare these VLMs for robot control. This paper revisits a practical yet seldom systematically studied question: which embodied VQA supervision actually transfers to downstream VLA policies? We introduce an automatic episode-to-VQA pipeline that turns simulated and real-world robot episodes into VQA, separating what a question teaches, which episodes it comes from, and how its answer is presented. Using this pipeline, we adapt a VLM with 16 generalist, capability-specific, and task-specific VQA recipes and train every adapted backbone under the same VLA protocol. Through closed-loop evaluation under diverse perturbations, we find that VLM benchmark averages are poor predictors of policy success. For instance, cross-view consistency VQA yields the strongest policy despite a lower VLM average than single-view spatial VQA. This challenges the common practice of selecting embodied supervision by VLM benchmarks. We further find that the gains are conditional. Cross-view supervision improves robustness to camera and especially robot-state changes, and much of the net improvement comes from preserving the policy's existing successes. Contrary to intuition, changing only the answer interface from multiple choice to direct generation raises the VLM average but lowers policy success. Finally, VQA built from real-world episodes also improves the policy, and mixing it with simulated VQA outperforms the simulation-only recipe at the same budget. These results indicate that embodied supervision should be chosen by the control conditions it improves and the successes it preserves.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.