acceptodds
Under review as a conference paper at ICLR 2027

When Target Accuracy Hides Semantic Interface Drift in Vision–Language Adaptation

Abstract

A target-adapted vision–language model can improve fine-grained accuracy while moving away from the pretrained image–text interface that supports zero-shot transfer. We define semantic interface drift through changes in direct prompt compatibility, prototype relations, local neighborhoods, and held-out zero-shot behavior. We study same-image frozen-feature anchoring in an identity-initialized visual adapter and test the role of the reference with a controlled intervention. The central experiment closely matches adaptation-validation accuracy while replacing the same-image reference with a within-minibatch shuffled feature. Across five-seed CLIP experiments, the same-image reference preserves the measured geometry and held-out transfer axes at comparable adaptation accuracy. The shuffled-reference control yields a 27.49-point mean transfer drop, compared with 3.78 points for the same-image reference. Three-seed replications recover the retention pattern with SigLIP2 and when EuroSAT replaces CIFAR-100 as the adaptation dataset. These results identify reference correspondence as a design axis that target accuracy alone does not measure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.