Mask-Guided Multi-Granularity Contrastive Learning for Fine-Grained Vision–Language Alignment
Abstract
Contrastive vision-language models such as CLIP excel at global image-text alignment but remain limited in fine-grained understanding. Existing methods construct local visual evidence by either implicitly selecting text-relevant patches or extracting region features with box-based operations such as RoIAlign. The former lacks explicit spatial grounding, while the latter often includes background or neighboring objects. To address these limitations, we introduce MaskAlign, a multi-granularity vision-language learning framework centered on pixel-level mask supervision. Masks provide precise spatial support, enabling the model to aggregate native visual patches predominantly covered by the target while suppressing background-dominated patches. We further construct MG-Data, comprising 6M images with long captions of over 100 words on average, short captions, and approximately 36M object and 14M region annotations, each pairing a curated mask with a detailed description. Using MG-Data, MaskAlign establishes image-, object-, and region-level visual-textual correspondences and integrates locally grounded semantics into global representations through cross-granularity alignment. We also introduce Long-caption Partial-Order (LPO), which promotes the use of additional details in long captions by enforcing a larger long-over-short similarity gain for matched pairs than for mined hard negatives. Extensive experiments demonstrate strong performance across long- and short-text retrieval, zero-shot classification, and fine-grained localization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.