acceptodds
Under review as a conference paper at ICLR 2027

Failure-Aware Offline Distillation for Efficient Vision-Language-Action Models

Abstract

We present an offline distillation pipeline that transfers a large vision-language-action (VLA) teacher into a compact student, with one recipe across teachers, student backbones, and task suites: (3.3B) and UniVLA (7B) teachers, Long-CLIP (159M) and SigLIP2 (1.15B) students, and the four LIBERO suites. The pipeline has three parts. First, the student trains on success trajectories that the teacher collects once. Second, rule code labels teacher frames with eight behaviour phases from gripper commands and physical events recovered by replay, and the same rule code together with a vision-language model (VLM) judge finds the phase in which each student failure occurs, with a third model resolving their disagreements (rule code + VLM judge). These phases match an author's blind check on 46 of 47 failures. Third, the student is retrained on the same teacher data with twice the loss on the phases that hold at least one third of the failures on its suite. Phase labelling and phase weighting bring the students close to their teachers. Success rises on all 14 weighted configurations (125 episodes rescued, 68 broken, ), and on LIBERO-10 the weighted Long-CLIP and SigLIP2 students reach 88.0% and 90.5%, against their UniVLA teacher's 91.5%. Training a weighted student, including its unweighted twin and the judge calls, costs 8.2 or 12.3 USD. At inference the students take 37 and 172 ms per forward pass on a 12 GB GPU that cannot hold UniVLA, against 431 ms for .

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.