acceptodds
Under review as a conference paper at ICLR 2027

Teachers Are Not Equal: Task-Flow Conditional Cross-Layer Knowledge Distillation for Efficient Flow-Based VLA Models

Abstract

Vision-Language-Action (VLA) models based on flow matching are expensive to deploy due to repeated Transformer computation at inference. One solution is to distill knowledge from a teacher VLA to a lightweight student VLA. Existing methods select a fixed teacher layer for distillation, but this may be suboptimal because layer informativeness varies with task instructions and flow stages, limiting student capacity. To address this, we introduce Task-Flow Conditional cross-layer Knowledge Distillation (TFC-KD), a framework that adaptively selects teacher layers conditioned on task instructions and flow stages. For each retained student layer, we first identify several candidate teacher layers around its initially aligned counterpart. We then use task instructions to modulate which teacher layers provide stronger supervision. Crucially, we divide the continuous flow process into three stages, allowing the preferred teacher layers to shift as action generation progresses. Task instructions and flow stages are combined to assign weights to candidate teacher layers, whose weighted attention maps over vision-language tokens serve as distillation targets. Finally, we impose an ordering constraint to keep teacher-layer selections roughly aligned with student depth while permitting local adaptation across task instructions and flow stages. On LIBERO, the trained 8-layer student achieves success, outperforming the Shallow- baseline by . When further compressed to 6 layers, TFC-KD incurs only a success drop relative to the teacher (), while reducing deployed parameters by and FLOPs by , with a speedup of inference on one A800 GPU.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.