Super Neurons Answer Before VLMs Speak
Abstract
Vector-level representations in pretrained vision-language models can be repurposed for prediction, but how small can these representations be? We show that direct scalar readouts from a handful of neural units (e.g. 10) in the LLMs of VLMs can outperform the full, billion-parameter model from which they are selected on categorical visual question-answering tasks. Notably, these super neurons can be identified in a gradient-free setting, using a single GPU. We thoroughly analyze super neurons and find that they satisfy key properties of effective predictors: they do not seem to exploit spurious biases and generalize well. Therefore, they enable diverse applications. Since they are discovered during prefill encoding, their predictions are able to bypass the autoregressive generation process of the LLM, dramatically reducing runtime by up to 27.8× on Qwen3-VL-8b-Instruct. Since they are sparse, they enable lightweight finetuning, preserving some of the VLM’s general capabilities while improving its performance on a target task. Our code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.