acceptodds
Under review as a conference paper at ICLR 2027

SemAE: Semantic-Aligned Action Expert for Language-Following Robot Control

Abstract

Action experts are crucial components for modeling robot actions, especially for transferring semantic knowledge from vision-language model (VLM) to high-dimensional, complex actions. However, recent studies have revealed that existing action experts can degrade the semantic knowledge contained in pretrained VLMs, leading to VLAs that, despite excelling at complex and dexterous manipulation, often fail to follow language instructions accurately. To address this challenge, we propose SemAE, an encoder-decoder action expert that enables mutual transformation between robot actions and semantic embeddings, allowing VLAs to seamlessly inherit the knowledge in VLMs. Specifically, we introduce motion primitive distillation on the action encoder, which enables it to compress robot actions into a compact set of latent embeddings that are well-aligned with the semantic space of the pretrained VLM. Moreover, we propose a patch-wise action flow decoder, which efficiently decompresses a single semantic embedding into future action sequence. Built on SemAE, SemVLA achieves strong language-following with fast inference via parallel decoding of semantic action embeddings. Extensive experiments across multiple standard benchmarks and real robots demonstrate that SemAE effectively aligns robot actions with semantic knowledge, substantially enhancing the language-following capabilities of Vision-Language-Action models. The source code will be publicly available on our project page

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.