MatryoshkaFusion: Multi-Granularity Visual Feature Fusion within Patch Regions for Multimodal Large Language Models
Abstract
Multi-granularity understanding is essential for visual comprehension, as the human visual system processes visual signals at multiple granularities. However, multi-granularity understanding mechanisms remain scarce in MLLMs, a mainstream paradigm in vision-language understanding. A limited number of studies have introduced multi-granularity understanding to accelerate inference at the cost of some performance degradation, leaving the potential for enhancing visual understanding largely unexplored. In this paper, we propose MatryoshkaFusion, a novel Matryoshka multi-granularity method that fuses multi-granularity features within patch regions, thus enhancing the model's visual understanding ability without increasing the number of LLM input tokens or the volume of training data. Furthermore, we visually demonstrate and analyze the weight distribution of different granularities, offering a new interpretable perspective on the mechanism underlying multi-granularity visual features’ enhancement of MLLMs’ visual understanding ability. To validate the effectiveness of our method, two sets of experiments are conducted: (1) training a specialized model on the ScienceQA dataset to rapidly assess the feasibility of our method; (2) training a multimodal chatbot and evaluating its zero-shot performance across multiple benchmarks. In both experiments, our model achieves better results than existing advanced multi-granularity methods used in MQT-LLAVA and LLaVA-M.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.