Meta-MemGran: Rethinking Pixel-Level Visual Encoder via Meta-Memory for MLLMs
Abstract
Recent progress in multimodal large language models (MLLMs) has largely depended on **CLIP**-based visual encoders. While these models excel at capturing global semantic alignment, they often struggle to preserve fine-grained visual details. By comparison, the recent **DINOv3** has exhibited emergent capabilities in fine-grained pixel-level perception, but still lacks the higher-level semantic abstraction required for alignment with LLMs and multi-granularity reasoning. These complementary limitations reveal a fundamental bottleneck: existing visual representations struggle to simultaneously preserve perceptual details and organize them into high-level semantics. We argue that the key challenge lies not merely in the visual encoder, but in the lack of an adaptive mechanism for transforming dense visual evidence into representations at different levels of abstraction. This raises a central question: ***How can fine-grained perception be effectively bridged with high-level semantics for multi-granularity multimodal reasoning?*** To address this challenge, we propose **Meta-MemGran**, a **DINOv3**-based MLLM with meta-memory-driven representation learning. By structuring and dynamically aggregating visual information through meta-memory, **Meta-MemGran** bridges pixel-level details, visual structures, and semantic abstractions without explicit granularity selection. Extensive experiments show that **Meta-MemGran** achieves an accuracy improvement of ** 40%↑** and a hallucination reduction of ** 30%↓**. Code is available at [Meta-MemGran](https://anonymous.4open.science/r/Meta-MemBank-BA1F/).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.