How native are native vision-language models?
Abstract
Vision-language models (VLMs) have been built from a vision encoder, a language model, and a connector between them, with the architecture prescribing how vision reaches language. Native VLMs remove these boundaries. How vision and language work together is learned rather than prescribed, yet removing the encoder and connector does not remove their work. The backbone must learn both to build visual features and to let language use them, and whether it does remains largely unexamined. Inside public native VLMs, we find that text draws on visual features mainly in the middle layers, even when they form much earlier. In controlled training, a backbone can learn visual features yet barely change its text predictions when the image is replaced, a dissociation reminiscent of optic aphasia, in which patients recognize objects but cannot name them. Through an external connector, the same features support far stronger image dependence, locating the bottleneck largely in the connection. We therefore introduce the adaptive visual reader, a connector learned inside the backbone, which raises caption precision from 20.7% to 37.0% while leaving visual features largely unchanged. The connector has not disappeared from native VLMs so much as become something they must learn. Their promise is a network that learns for itself how vision and language work together, and realizing it calls for training and evaluation that treat visual encoding and reading as distinct problems and develop both.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.