acceptodds
Under review as a conference paper at ICLR 2027

feiner: Preserve, Project, and Let the Decoder Integrate Visual Information

Abstract

Extending language models to reason over additional modalities is an important step towards general-purpose AI models. Since not all modalities can be efficiently trained natively, integrating multimodal information into pre-trained language decoders is an important task. In this work, we focus on vision, one of the most important non-text modalities, and use LLaVA-style vision-language models to study how information from vision encoders can be preserved and utilized in pre-trained decoders. To this end, we introduce a framework for efficient, controlled training and evaluation. Our resulting design preserves strong modality-specific representations and effectively projects them into the decoder's latent space to increase the availability of vision-specific information. We find that representation preservation and decoder-side information flow are complementary requirements for strong visual understanding. Our *feiner* approach rivals state-of-the-art models at comparable scale while requiring less than half the compute. These results provide an efficient basis for studying how pre-trained decoders can incorporate additional modalities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.