Unbundling the Vision Encoder in Multimodal LLMs
Abstract
How multimodal LLMs should process images remains unsettled. Most rely on a vision encoder optimized separately for its own objective, while recent recipes train the encoder from scratch alongside the language model or remove it entirely. We view the vision encoder as a bundle of three choices that these recipes keep or discard: a visual prior from separate pretraining, dedicated vision-only capacity, and placement of that capacity before the multimodal backbone (pre-fusion) rather than inside it (in-fusion). Because published models also differ in data, backbone, and compute, we test each choice in a controlled study of 1.7B–8B multimodal LLMs trained from scratch on common data, evaluated on downstream visual question answering and language benchmarks at matched lifetime compute. An encoder trained from scratch with the backbone matches pretrained ones, so separate pretraining offers little advantage once its compute is counted. What matters instead is dedicated capacity: encoder-free models trail on visual understanding even with the encoder's parameters added to their backbone. Yet that capacity need not precede fusion, as vision-only modules in early backbone layers (in-fusion) match or slightly exceed the encoder given longer training, while lowering time to first token. For multimodal LLMs trained from scratch, our findings favor a single training stage with dedicated vision capacity early in the network.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.