acceptodds
Under review as a conference paper at ICLR 2027

Heard but Not Heeded: Language Survives in a VLA's Representations but Not Its Actions

Abstract

A fine-tuned Vision-Language-Action (VLA) model that scores on the LIBERO manipulation benchmark stops following language instructions when given a scene outside its training distribution. We diagnose this failure through three measurements: (1) the model still internally represents which object was named ( accuracy vs baseline), (2) attention roughly doubles when the object is mentioned, yet (3) the policy grasps the wrong object. This dissociation understanding language and attending correctly, but choosing the wrong action shows that instruction-following failures need not stem from language misunderstanding or visual misidentification. We validate this dissociation with a within-distribution control experiment. Attention allocation to the target object is virtually identical in both settings ( vs of total attention), yet language-driven attention shifts are five times stronger when behavior fails. This suggests that insufficient attention magnitude cannot account for the behavioral failure. The evidence points to a breakdown at the interface where the model translates visual and language information into motor commands. We introduce object-level behavior metrics (recording which object was actually grasped, rather than only whether the task succeeded) and show that benchmark structural properties make instruction-following failures invisible to standard evaluation metrics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.