acceptodds
Under review as a conference paper at ICLR 2027

ReMo: Rethinking Modality Roles in Vision-Language Models for General Visual Perception and Reasoning

Abstract

Despite remarkable strides in multimodal understanding, the general visual perception of Vision-Language Models (VLMs) remains a critical weakness. Existing solutions typically rely on resource-heavy knowledge distillation or latency-intensive external visual tools, neither of which fundamentally improves the VLM's intrinsic perceptual mechanisms. In this paper, we present , a novel framework that thinks the complementary dality roles of language and vision to achieve intrinsic visual enhancement in VLMs. Our core insight is that language should actively orchestrate visual perception rather than merely respond to it. ReMo first leverages the expansive reasoning capabilities of Large Language Models (LLMs) to generate linguistic priors, dynamically extracting reasoning-aware guidance from cross-modal attention to localize task-relevant visual evidence. However, since cross-modal alignment often discards fine-grained visual details, we introduce a structural complementary mechanism. By exploiting the innate relationship modeling of Vision Transformers (ViTs), ReMo actively recovers vital visual contexts that are typically overlooked by the LLM. By jointly harmonizing language-guided visual search and ViT-driven visual recovery alongside visual evidence reinforcement, our framework significantly elevates the general visual perception and reasoning capacities of VLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.