Concepts Are Distributions: Training-Free Insertion in Vision-Language Models
Abstract
Many activation-steering methods represent a concept as a direction in activation space. We test whether this representation also supports object insertion in the visual tokens of frozen vision-language models (VLMs). We find a clear asymmetry. A single direction can suppress an object, but it does not reliably introduce an object that is absent from the image. Learned directions, text-derived directions, and individual real activations all perform poorly. We show that the limitation is distributional. Repeating one concept vector across visual tokens can match the mean while collapsing covariance. A Gaussian whose population mean and covariance match those of real object activations produces responses comparable to transplanting a real object, whereas a surrogate that preserves the marginal statistics but destroys covariance does not. Existing feature-transfer operators follow the same ordering as the second-order structure they transfer. We therefore store each concept as a low-rank Gaussian estimated from a few images. On CLIP-L/14, four covariance directions are sufficient, giving a 20 KB representation per concept. Other vision encoders require modestly higher rank. The method requires no training and no reference image at inference time, transfers across five models and three vision encoders, and works without object-location annotations. These results show that object insertion depends on second-order visual-token structure and identify the visual-to-language activation stream as a security-sensitive interface when internal activations can be modified.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.