acceptodds
Under review as a conference paper at ICLR 2027

What Behavior Cloning Gives and Takes: Action Learning and Visual Reorganization in VLAs

Abstract

A VLA can forget an answer without forgetting what it saw. During adaptation on LIBERO, a base-fitted localization readout degrades from 1.55 to 6.28 cm, yet refitting the same readout family on the adapted checkpoint reduces error to 1.13 cm. The visual information remains linearly recoverable; the pre-adaptation reader no longer matches its coordinates. We follow two behavior-cloning runs and align closed-loop control with native heads, fixed and refit probes, six-site representations, and matched interventions. The trajectory has three stages. Pre-adaptation readouts fail by 500 updates while success is still 0–5%. Control then rises to 85–88% by 4–8k while spatial information remains recoverable. As success reaches 94–97%, part of the checkpoint-local recoverability declines. A rank-64 orthogonal alignment closes over 95% of the destination-readout gap, the native language head fails on a separate schedule, and the action expert builds another destination code behind a frozen backbone. Gradient conflict arrives after the first remapping, and freezing its strongest layers offers no advantage over freezing random layers. Freezing and replay preserve selected readouts at an action cost; an auxiliary action-token objective accelerates the earliest action endpoint. Behavior cloning reorganizes the relations among representations, readers, and actions. Training should maintain or deliberately relearn those relations instead of preserving parameters indiscriminately.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.