FAKEVLA: A PROBE FOR ANALYZING VISUAL REPRESENTATIONS FOR ROBOT ACTIONS
Abstract
Vision encoders are commonly evaluated frozen through downstream probes that reveal what information is encoded and how readily it can be used for prediction. In robotic manipulation, however, visual representations are typically evaluated within complete visuomotor policies, making it difficult to separate representation quality from downstream control. We propose FakeVLA, a framework for probing frozen visual representations through manipulation-specific tests that connect feature geometry to action selection. FakeVLA uses nearest-neighbor (NN) retrieval over demonstrations to test whether local visual neighborhoods recover manipulation-relevant physical states, finding that stronger state recovery is closely associated with higher Grasp and Place success. Further lightweight action predictors are used to test whether useful information remains decodable beyond local NN geometry. We evaluate generic, vision-language, and VLAderived representations under controlled state and visual shifts, and compare probe trends with corresponding full-VLA policies whose vision modules remain frozen. Lower Drift is strongly associated with higher manipulation success, while texture causes the largest robustness degradation. Nonlinear predictors perform strongly under familiar observations but degrade more sharply under visual shifts than NN retrieval, showing that action decodability and local geometry capture distinct properties. Probe performance is also positively associated with full-VLA performance. Finally, strong within-domain performance does not imply sim-to-real alignment: cross-domain retrieval remains substantially weaker, although VLA pretraining and real-domain adaptation improve SimReal transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.