acceptodds
Under review as a conference paper at ICLR 2027

Mask-Guided Multi-Granularity Contrastive Learning for Fine-Grained Vision–Language Alignment

Abstract

Contrastive vision-language models such as CLIP excel at global image-text alignment but remain limited in fine-grained understanding. Existing methods construct local visual evidence by either implicitly selecting text-relevant patches or extracting region features with box-based operations such as RoIAlign. The former lacks explicit spatial grounding, while the latter often includes background or neighboring objects. To address these limitations, we introduce MaskAlign, a multi-granularity vision-language learning framework centered on pixel-level mask supervision. Masks provide precise spatial support, enabling the model to aggregate native visual patches predominantly covered by the target while suppressing background-dominated patches. We further construct MG-Data, comprising 6M images with long captions of over 100 words on average, short captions, and approximately 36M object and 14M region annotations, each pairing a curated mask with a detailed description. Using MG-Data, MaskAlign establishes image-, object-, and region-level visual-textual correspondences and integrates locally grounded semantics into global representations through cross-granularity alignment. We also introduce Long-caption Partial-Order (LPO), which promotes the use of additional details in long captions by enforcing a larger long-over-short similarity gain for matched pairs than for mined hard negatives. Extensive experiments demonstrate strong performance across long- and short-text retrieval, zero-shot classification, and fine-grained localization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.