Adapters Do More Than Adapt: Coupling Parameter-Efficient Fine-Tuning with Recovery-Aware Structured Pruning for Visual Grounding
Abstract
Visual grounding aims to localize an image region specified by a natural-language expression. Large pretrained vision-language models have substantially improved grounding accuracy but incur considerable adaptation and inference costs. Parameter-efficient fine-tuning reduces trainable parameters while leaving the deployed backbone intact, whereas structured pruning produces smaller models but is typically decoupled from task adaptation, leading to representation shifts and degraded task-relevant responses. To bridge this gap, we propose Adapter-Conditioned Recovery Pruning (ACRP), which couples parameter-efficient adaptation with recovery-aware structured pruning in a unified optimization framework. The key idea is to counterfactually supervise adapters to approximate compact signatures of candidate-removed responses, thereby providing recovery-aware guidance for pruning. ACRP combines modality-aware joint budget allocation across attention and MLP structures with a three-stage optimization strategy, allowing adaptation and compression to interact before physical pruning. Moreover, parallel multi-level feature branches integrate hierarchical visual semantics, while an evidence-preserving token condenser retains expression-relevant tokens and aggregates the remaining tokens into regional context. Experiments on four public datasets demonstrate that ACRP achieves effective task adaptation and retains competitive localization performance, while updating merely 0.38% of the pre-trained encoder parameters and physically pruning approximately 40% of the backbone parameters.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.