Knowing Is Not Acting in VLA Models
Abstract
Vision-language-action (VLA) policies can identify an object from an indirect description yet fail to select it when acting. We study this knowledge-action gap with KnowAct, which pairs literal and knowledge-based references to the target and an alternative object in balanced manipulation scenes. Its contrastive knowledge utilization (KU) measures knowledge-conditioned selection relative to literal control on a fixed set screened for reported knowledge and initial control. We also introduce Referential Self-Distillation (RSD), which trains direct execution against an EMA teacher conditioned on a translated object name. We evaluate GR00T N1.7 and in RoboTwin 2.0. On GR00T, RSD reaches KU 0.311 versus 0.210 for a collection-protocol-matched DAgger control; the paired item-bootstrap difference is 0.101 [0.067,0.135]. On , RSD reaches 0.345 versus 0.258 for DAgger, a paired difference of 0.087 [0.048,0.126]. Among 157 GR00T items, knowledge contrast improves on 110, worsens on 35, and is unchanged on 12. The effect is positive in each evaluated split; knowledge injection, readout probes, activation patching, and control experiments further delimit the finding. The item-bootstrap intervals condition on three trained runs and do not quantify retraining uncertainty.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.