acceptodds
Under review as a conference paper at ICLR 2027

Uncovering Separable and Composable Skill Mechanisms in Vision-Language Models

Abstract

Recent work has sought to localize the internal mechanisms underlying visual reasoning skills in vision-language models. We posit two important desiderata, separability and composability, for skill localization. Separability requires identified mechanisms to contribute primarily to their target skill while having limited effects on other skills, enabling precise diagnosis and targeted intervention. Composability requires skill mechanisms to be combinable to support complex tasks, allowing a small set of localized skill mechanisms to account for broader capabilities. However, existing methods overlook separability and composability in both mechanism identification and evaluation. To address these limitations, we propose a multi-skill interpretability framework DISentangled and COmposable Skill Mechanisms (DISCO) for identifying and evaluating separable and composable attention heads. For mechanism identification, we propose a novel technique called Separable Head Selection, which selects skill-specific heads based on their contributions across skills. For evaluation, we are the first to design metrics for evaluating separability and composability of identified heads. Experiments on three VLMs and four fundamental visual reasoning skills show that the heads identified by DISCO achieve the highest overall separability on single-skill tasks, outperforming the strongest baseline by 22.75, 10.95, and 9.96 points across the three models, while also achieving the highest average composability on multi-skill tasks across all models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.