Uncovering Separable and Composable Skill Mechanisms in Vision-Language Models
Abstract
Recent work has sought to localize the internal mechanisms underlying visual reasoning skills in vision-language models. We posit two important desiderata, separability and composability, for skill localization. Separability requires identified mechanisms to contribute primarily to their target skill while having limited effects on other skills, enabling precise diagnosis and targeted intervention. Composability requires skill mechanisms to be combinable to support complex tasks, allowing a small set of localized skill mechanisms to account for broader capabilities. However, existing methods overlook separability and composability in both mechanism identification and evaluation. To address these limitations, we propose a multi-skill interpretability framework DISentangled and COmposable Skill Mechanisms (DISCO) for identifying and evaluating separable and composable attention heads. For mechanism identification, we propose a novel technique called Separable Head Selection, which selects skill-specific heads based on their contributions across skills. For evaluation, we are the first to design metrics for evaluating separability and composability of identified heads. Experiments on three VLMs and four fundamental visual reasoning skills show that the heads identified by DISCO achieve the highest overall separability on single-skill tasks, outperforming the strongest baseline by 22.75, 10.95, and 9.96 points across the three models, while also achieving the highest average composability on multi-skill tasks across all models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.