acceptodds
Under review as a conference paper at ICLR 2027

ReGDec: Decoupling Region Cues from Global Semantics for VLM-Grounded Open-Vocabulary Dense Perception

Abstract

Vision-Language Pretrained Models (VLMs) achieve strong open-vocabulary perception by aligning visual representations with natural language semantics, but their global image-text pre-training objectives often limit local discriminability for dense perception. Existing approaches improve dense perception through feature refinement, external spatial guidance, or task-specific adaptation. In this paper, we analyze how spatial visual cues evolve across VLM depths and observe that preferred feature depths vary across tasks and regions. These observations motivate region-adaptive fusion of semantic and region cues across layers. To address this, we propose ReGDec (Region-Global Decoupling), a region-adaptive framework that decomposes region features into CLS-aligned components and orthogonal residuals. A compact selector learns separate region-conditioned weights to combine the two components across layers, while keeping the VLM encoders frozen. Experiments across multiple VLM backbones and dense perception benchmarks demonstrate the effectiveness of ReGDec, with gains varying across backbones and tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.