From Fish to Fishing: Teaching CLIP to Generalize through Circuit-Level Distillation
Abstract
CLIP aligns image content with text-defined categories, enabling dense image-text prediction. But generalizing its ability to unseen domains remains challenging. Feature-level distillation improves local semantics and spatial coherence, yet the resulting representations can still deteriorate under appearance shifts. We investigate this gap at both the representation and circuit levels, observing that the self-supervised vision foundation model (VFM) maintains more consistent dense features and greater overlap in prediction-relevant neurons than CLIP across the examined domain shifts. These observations motivate transferring not only region representations, but also the ability to reuse internal channel support across appearances. Based on that, we propose ReoNa, a novel framework that uses VFM-guided Region-to-channel distillation to teach CLIP to reuse the internal MLP channels supporting each semantic region across appearances. Extensive experiments across diverse open vocabulary domain generalized semantic segmentation (OVDG-SS) benchmarks demonstrate that ReoNa significantly improves CLIP dense representation despite domain shifts. Our code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.