Rosetta Neurons Across Vision and Language
Abstract
Can independently trained vision and language models develop individual neurons that respond to the same concepts? We identify these cross-modal rosetta neurons: pairs of neurons in separate vision and language models whose activations correspond across images and text. We extend neuron matching to paired images and captions by correlating each neuron's mean activation over the corresponding inputs. We find such matches across CLIP, DINOv2, MAE, and a Stable-Diffusion UNet paired with language models, including vision encoders trained without text. We recover the same neuron pairs in two separately sampled sets of image–caption pairs. Blinded model judges and human raters support concept agreement among selected high-correlation pairs. The matched population grows with model size, following sublinear power-law scaling across language models and text-free vision encoders. These findings suggest that independently trained models can represent corresponding visual and linguistic concepts at the level of individual neurons. We use these correspondences for text-indexed image retrieval: a concept named in text selects a language neuron, whose matched vision neuron retrieves images of that concept without additional training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.