Grounding Is Inherited, Action Is Concentrated: What VLA Finetuning Adds to a Web-Pretrained VLM
Abstract
A web-pretrained vision-language embedding model with a one-hidden-layer action head, adapted with LoRA on the target benchmark's demonstrations, reaches 4.62±0.05 of 5 consecutive subtasks on CALVIN ABC→D and 98.0% on the 40-task LIBERO generalist. We ask what the decodability of the action from the backbone's pooled feature tells us about the resulting policy, and what it does not. On LIBERO, a ridge reader recovers the action chunk from the frozen backbone's pooled feature only with labelled frames ( 0.64–0.68) and almost not at all from its eight leading directions; after finetuning the same reader reaches 0.95 from 400 frames and 0.90–0.93 from the top eight, across backbones, poolings and adaptation methods (CALVIN: the same trend at smaller magnitude). Working policies depend on those directions: keeping only the top 64 of 2,048–4,096 retains 88–96% success and 64 random directions retain none. Decodability does not certify reliability, however. Controlled input interventions on four full-finetuning policies, three of them probed and readable, switch whole task families between success and failure: on CALVIN, replacing only the table's wood texture with the geometry held fixed turns two skills off and on (223–226/226 ↔ 0/226 first subtasks, two backbones) while a LoRA policy is unaffected in the three combinations tested; on LIBERO, a one-token change of the evaluation prompt drops two policies from 94.8 to 46.9% and from 74.9 to 15.0% while top-8 demonstration-frame readability moves by 0.02–0.05 and no working policy tested changes comparably. Other deficits persist under every input tested. Across the configurations studied, finetuning increases action decodability and concentrates it in leading directions; these measurements by themselves do not establish closed-loop reliability. Simulation only; one arm.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.