Carve VFM: Carving Specialists out of Generalists, Automatic Task-Specific Compression of Vision Foundation Models
Abstract
Vision foundation models (VFMs) provide transferable visual representations, but their inference cost limits deployment in many settings. Fine-tuning specializes a VFM without shrinking it; distillation shrinks it, but through expensive teacher-guided training. Structured pruning offers a more direct route by extracting a compact subnetwork from the VFM. However, as compression gets aggressive, existing methods fail because they treat subnetwork selection and recovery as separate axes. Across layers, capacity must follow task sensitivity rather than spread uniformly. Within a layer, a removal cannot be judged by importance alone; its impact also depends on how well the lost contribution can be recovered from what remains. Selection and recovery must therefore be decided jointly. We introduce Carve VFM, which resolves this interdependence. Given a VFM, a calibration set, and a parameter budget, it (i) searches for task-specific capacity allocations across layers; (ii) guides the search via an adaptive recovery-aware fitness criterion and (iii) reconstructs removed behavior through iterative closed-form low-rank corrections merged into the surviving weights. The result is a dense task model produced in minutes on a single GPU, without training. Across twelve benchmarks and four models from DINOv3 and MetaCLIP-2, Carve VFM ranks first among all evaluated structured-pruning methods at every budget, averaging +9.2%; crucially, this margin widens from +3.0% at 35% parameter removal to +16.3% at 70%. Further fine-tuning surpasses the strongest evaluated edge model 88.1% vs. 83.4%, and carved models cut compute and memory while delivering 1.3-2.0x higher GPU throughput than the uncompressed one. These results show that the compact model you need is inside the VFM; it only has to be carved out.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.