acceptodds
Under review as a conference paper at ICLR 2027

Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving

Abstract

Autonomous driving is an important and safety-critical task, and recent advances in LLMs/VLMs have opened new possibilities for reasoning and planning in this domain. However, large models demand substantial GPU memory and exhibit high inference latency, while conventional supervised fine-tuning (SFT) often struggles to bridge the capability gaps of small models. To address these limitations, we propose Drive-KD, a framework that decomposes autonomous driving into a “perception–reasoning–planning” triad and transfers these capabilities via knowledge distillation to improve VLM performance on comprehensive driving tasks. Through a systematic study of distillation design for autonomous driving, we identify layer-specific attention as the distillation signal to construct capability-specific single-teacher models. Moreover, we unify these single-teacher settings into a multi-teacher distillation framework and introduce asymmetric gradient projection to mitigate cross-capability gradient conflicts. Experiments show that our distilled InternVL3-1B model, with less GPU memory and higher throughput, achieves better overall performance than the pretrained 78B model from the same family on DriveBench, and surpasses GPT-5.1 on the planning dimension. It further substantially outperforms its SFT counterpart in closed-loop evaluation on Bench2Drive-VL, achieving 84.50 Route Completion. Further attention analysis shows that Drive-KD restores the capability-specific visual grounding suppressed by SFT.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.