acceptodds
Under review as a conference paper at ICLR 2027

DARK: Diagonal-Anchored Repulsive Knowledge Distillation for Vision-Language Models under Extreme Compression

Abstract

Large vision-language models provide useful representations for fetal ultrasound, but their computational cost makes deployment on portable devices difficult. Knowledge distillation offers a route to smaller models, although continued imitation of a large teacher may constrain a much smaller student. We propose Diagonal-Anchored Repulsive Knowledge Distillation (DARK), which preserves supervision on matched image–caption pairs while gradually changing supervision on other pairs from imitation to repulsion. This lets the student first learn the teacher's relationships and then depart from them without abandoning matched-pair alignment. We use DARK to compress FetalCLIP into MobileFetalCLIP, whose visual encoder has approximately 26× fewer parameters and runs in 1.6 ms on an iPhone 16 Pro. Compared with static distillation, DARK improves brain sub-plane macro-F1 from 0.701 to 0.775. On an external cohort of 403 images from five African countries, anatomical macro-F1 rises from 0.770 to 0.895, with gains in all four classes. Ablations show that the sign change is essential: fixed decoupled weights, as in decoupled KD, and stopping imitation at zero both fall short. Linear probes on the compact model's frozen features retain 97–98% of the teacher's downstream performance. These results support selective departure from teacher imitation as a way to learn compact fetal-ultrasound representations for both zero-shot recognition and supervised downstream use.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.