acceptodds
Under review as a conference paper at ICLR 2027

Deep and Compact Visual Experts for Fast Encoder-Free MLLM

Abstract

Encoder-free Multimodal Large Language Models (MLLMs) process images and text within a single backbone, enabling early cross-modal fusion and eliminating the prefill latency incurred by vision encoders in modular MLLMs like LLaVA. This prefill advantage is particularly important for edge deployment and real-time interactions. Despite this advantage, most encoder-free architectures ignore the latency aspect. In this paper, we first examine the design space of encoder-free architectures to measure both multimodal performance and latency. Our analysis shows that existing encoder-free designs are limited by insufficient visual processing depth in shallow backbones. Building on this finding, we propose DeCoVE, a Deep and Compact Visual Expert architecture that decouples visual processing depth from backbone depth by trading visual width for depth. We then demonstrate the performance–latency tradeoff of DeCoVE by varying its depth. Moreover, DeCoVE can be initialized from parts of pretrained vision encoders, bringing together the training efficiency of pretrained models and the prefill advantage of encoder-free designs. Based on DeCoVE, we develop DeCoVE-3B, an encoder-free MLLM that achieves state-of-the-art performance among encoder-free baselines and is competitive with Qwen3.5-VL 2B, while reducing prefill latency by 41.5% across 13 benchmarks. DeCoVE-3B establishes a new Pareto frontier between performance and prefill latency for MLLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.