acceptodds
Under review as a conference paper at ICLR 2027

Depth-wise Specialization in Vision-Language Models for Compositional Adaptation

Abstract

Contrastive vision–language models such as CLIP and SigLIP achieve strong global image–text alignment, yet often struggle with fine-grained compositional distinctions such as modifier sensitivity and visually grounded relations. Prior work improves compositionality through stronger supervision, caption enrichment, or pooling modifications, but largely overlooks how compositional cues evolve across encoder depth. We reveal complementary depth specialization: noun-phrase selectivity strengthens toward later layers, whereas modifier- and relation-sensitive distinctions progressively attenuate. Depth-wise probes of broad noun-phrase selectivity, same-head modifier sensitivity, and relation-family discrimination consistently support this pattern. Motivated by this finding, we introduce RELIP, a depth-asymmetric fine-tuning method for contrastive dual encoders. RELIP retains each backbone's native global image–caption objective while using shallow patch representations as compositional evidence and a detached late relevance map to emphasize attribute–entity and grounded-relation cues that attenuate with depth. The auxiliary branch is training-only and leaves the original inference architecture unchanged. Under the same CC300K budget, RELIP improves both CLIP and SigLIP; with SigLIP, it reaches on SugarCrepe and on SugarCrepe++, outperforming CLIP by and points while preserving strong retrieval and zero-shot performance. These results show that effective compositional adaptation depends not only on what supervision is provided, but also on where and how it is applied.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.