acceptodds
Under review as a conference paper at ICLR 2027

SmartVLA: A Steerable Action Interface for Fine-Grained, Multimodal, and Corrective Instructions

Abstract

Existing vision-language-action (VLA) models can perform diverse manipulation tasks from task-level instructions, but strong task performance does not ensure fine-grained semantic following, multimodal instruction understanding, or mid- execution corrective control. To this end, we propose SmartVLA, which con- structs semantic and spatial supervision through action-aligned segmentation and grade-conditioned instruction augmentation, and combines warmup, pretraining, and adaptation to explicitly learn the mapping from current instructions to local actions. The resulting unified instruction interface supports natural human steering and enables SmartVLA to serve as a low-level control interface for higher-level agentic systems. To systematically assess these three capabilities, we introduce LIBERO-SMART, a benchmark comprising Semantic, Visual Prompt, and Con- trol suites with 120 tasks and 6,000 evaluation instances. We also evaluate basic task performance, generalization across distributions, and real-world deployment on LIBERO, LIBERO-Plus, LIBERO-Pro, RoboTwin, and the Piper real-robot platform. Experiments show that SmartVLA preserves strong task performance while more reliably following diverse instructions, visually specified commands, and instructions updated during execution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.