Editing Visual Capabilities of Large Vision-Language Models through Task Arithmetic
Abstract
Large vision-language models (LVLMs) are typically adapted through end-to-end multimodal training, even when the desired modification primarily concerns visual capabilities. In this work, we study whether such capabilities can be upgraded independently from the rest of the model, as is common in human-engineered modular systems. We show that LVLMs support localized interventions on the vision tower: task updates can be constructed from independently specialized CLIP vision encoders and applied directly to the LVLM vision tower, while keeping the multimodal adapter and language model fixed. This provides an efficient means of transferring visual expertise into LVLMs without additional multimodal training and, when available, enables the reuse of fine-tuned or domain-specialized CLIP checkpoints. We demonstrate that this intervention supports multiple editing goals, including improving a single downstream task, composing improvements across multiple tasks, and removing targeted visual capabilities while preserving model behavior. Our results show that vision-encoder editing improves target-task performance while maintaining the original multimodal capabilities of the LVLM, at a fraction of the cost of end-to-end training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.