Post-Connector Distillation: An Efficient Recipe for Training Vision Encoders
Abstract
Different efficient vision encoder architectures are needed for different hardware, but developing a new vision encoder requires a large amount of compute, elaborate training pipelines, and repeated downstream fine-tuning to measure its quality on VLM tasks. For fast iteration, we introduce post-connector distillation, a simple two-stage recipe for learning from an existing VLM. The first stage is a simple student-teacher reconstruction loss over the final multimodal connector output embeddings, which requires only unlabeled images or videos. This stage enables direct evaluation on downstream VQA benchmarks without additional instruction tuning, significantly accelerating model development. The second stage jointly refines the encoder, connector, and language backbone through supervised fine-tuning (SFT). It closes the remaining capability gaps between the student vision encoder and teacher vision encoder. In our ablations, training loss is a poor indicator of encoder quality, while the early benchmark scores predict which data mixture and architecture perform better after fine-tuning. Paired with a different language backbone, it matches Qwen3.5-2B's ViT across 13 VLM benchmarks while requiring 2.2 fewer encoder-plus-connector FLOPs, and outperforms all other evaluated encoders.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.