acceptodds
Under review as a conference paper at ICLR 2027

LANCE: Learning What to See in Vision-Language-Action Policies via Language-Anchored Action Queries

Abstract

Vision-language-action (VLA) policies have advanced robotic manipulation, yet they still struggle with reliable instruction following. These policies tend to exploit vision-action correlations instead of consistently using language to identify task-relevant visual evidence for action prediction. To address this issue, we propose **LANCE** (**Language-Anchored Action Queries**), an action-query framework that explicitly couples scene-conditioned action representations to task semantics derived from language instructions. LANCE employs two sets of learnable action queries: one reads only the language tokens to construct a *language anchor*, while the other reads both language and visual tokens to form a scene-conditioned action representation. To explicitly couple the scene-conditioned representation with the language anchor, *Reciprocal Alignment* aligns their latent distributions. Crucially, alignment alone does not ensure that the language anchor captures the action distinctions between instructions. We therefore devise *PriorAction*, an auxiliary action prediction loss conditioned on the language anchor. The two objectives jointly steer visual processing to prioritize task-relevant evidence for instruction-consistent action prediction. Experiments across diverse unseen settings show that LANCE improves instruction-consistent target selection. Moreover, under zero-shot composition of learned skills, LANCE achieves higher semantic coverage and subtask completion rates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.