acceptodds
Under review as a conference paper at ICLR 2027

Learning to Route Patches: Mixture of Patch Embeddings Vision Tokenization for MLLMs

Abstract

Multimodal large language models (MLLMs) traditionally rely on a fixed visual tokenization scheme, where all images are discretized using a single patch size, regardless of input complexity or task requirements. This design enforces a uniform trade-off between representation fidelity and computational cost, which we find to be suboptimal. In this work, we investigate the role of visual token granularity under a fixed token budget. Our results show that reducing patch size consistently yields larger performance gains than increasing image resolution at comparable computational cost, indicating that tokenization granularity is a more effective factor in improving multimodal reasoning. However, finer tokenization is not universally optimal, and different benchmarks exhibit distinct performance saturation points, suggesting that the best patch size varies depending on the input and task. Motivated by this, we propose Mixture of Patch Embeddings (MoP), an adaptive visual tokenization framework that jointly trains multiple patch embedding modules and learns a lightweight routing policy to select the most suitable patch size for each input. MoP offers instance-wise control over representation granularity, enabling flexible performance optimization. Extensive experiments across diverse multimodal benchmarks demonstrate that MoP consistently outperforms fixed patch size baselines, delivering the best performance on 7 benchmarks and reducing FLOPs by over 1.4× compared to fixed patch tokenization. Our results highlight the potential of adaptive tokenization in advancing MLLMs, where the flexibility of patch size selection enhances model performance without necessitating significant efficiency trade-offs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.