acceptodds
Under review as a conference paper at ICLR 2027

MatryoshkaFusion: Multi-Granularity Visual Feature Fusion within Patch Regions for Multimodal Large Language Models

Abstract

Multi-granularity understanding is essential for visual comprehension, as the human visual system processes visual signals at multiple granularities. However, multi-granularity understanding mechanisms remain scarce in MLLMs, a mainstream paradigm in vision-language understanding. A limited number of studies have introduced multi-granularity understanding to accelerate inference at the cost of some performance degradation, leaving the potential for enhancing visual understanding largely unexplored. In this paper, we propose MatryoshkaFusion, a novel Matryoshka multi-granularity method that fuses multi-granularity features within patch regions, thus enhancing the model's visual understanding ability without increasing the number of LLM input tokens or the volume of training data. Furthermore, we visually demonstrate and analyze the weight distribution of different granularities, offering a new interpretable perspective on the mechanism underlying multi-granularity visual features’ enhancement of MLLMs’ visual understanding ability. To validate the effectiveness of our method, two sets of experiments are conducted: (1) training a specialized model on the ScienceQA dataset to rapidly assess the feasibility of our method; (2) training a multimodal chatbot and evaluating its zero-shot performance across multiple benchmarks. In both experiments, our model achieves better results than existing advanced multi-granularity methods used in MQT-LLAVA and LLaVA-M.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.