acceptodds
Under review as a conference paper at ICLR 2027

Channel-Remix: Data efficient Synthesis of Task-Specific Compact Models by Remixing Vision Foundation Model Channels

Abstract

Large-scale pretraining endows vision foundation models (VFMs) with rich, transferable visual knowledge, yet their compute and memory costs put much of it out of reach in practice. Knowledge distillation compresses this knowledge by transferring it to a compact student under the VFM's supervision. However, while such supervision provides a strong training signal, it leaves the student's parameter space unconstrained. Consequently, bridging the gap between teacher representations and viable student weights typically demands substantial task data and long training. We propose Channel-Remix, which uses the VFM not only as a teacher but as a library of pretrained weight directions from which the student is built. This recasts knowledge transfer as selecting and recombining channels that pretraining has already learned, a problem that requires far less task data than learning new weights. Given a parameter budget and a small calibration set, Channel-Remix retains a core subset of channels and sparsely mixes pruned ones into them to form new task-adapted channels with zero runtime cost. By grounding both the learning signal and the student's parameter space in the VFM, Channel-Remix achieves high data efficiency, matching LoRA-based distillation using 3-4x less calibration data. Across 16 tasks, 7 pretraining families, and 5 model sizes, it improves pruned models by 4.6% on average and closes the gap to the original VFM to under 2% at 50% compression, outperforming distillation baselines under full fine-tuning, PEFT, and post-pruning reconstruction while training fewer parameters. Moreover, because every student is built from the same pool of VFM directions, models for different tasks become directly comparable, opening a route to analyzing which knowledge is shared across tasks and which is task-specific.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.