Visual Instruction Channels for Generalist Robot Policies
Abstract
Vision-Language-Action (VLA) models offer a capable foundation for robot control, but their usefulness depends on the ability to accurately steer them, that is to specify what a robot should do. Natural language provides a flexible, yet sometimes ambiguous and inaccurate, interface, further limited by the reliability of language-steering. We perform an in-depth analysis of visual instructions as conditioning, individual and complementary to language, supplying them through the policy's existing image inputs without architectural modification. In simulation, we vary the signal's shape (point, box, mask), its delivery format (overlay, blur, blackout, separate image), and its quality (view coverage and signal delay). We find that quality matters most: a signal updated at every observation substantially outperforms the static, first-frame signal often used in prior work. Shape and format effects depend on the backbone and task, with masks and boxes providing reliable overlay choices across the evaluated policies. On three real-robot tasks, visual instructions improve task performance over language alone in every case. These results highlight visual instructions as a practical channel for communicating task intent to generalist robot policies, one that can outperform language on its own and is strongest in combination with it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.