acceptodds
Under review as a conference paper at ICLR 2027

Coordinates or Equivalence? Alignment Reallocates Sensitivity Across Input Variations

Abstract

Contrastive image-text training changes how an encoder responds to variations that its captions may leave unspecified. We measure this change using distributions induced by a fixed panel of input transformations. Each family's within-input covariance is compared with the covariance of the complete mixture, yielding response shares and cross-family covariance similarities that are invariant to invertible affine reparameterisations. An exact covariance budget separates within-component variation from variation of component means. A conditional linear-model analysis further expresses response changes through the encoder's input-space readout operator. On Stanford Cars, disrupting image-caption correspondence changes the mean per-seed correlation between a frozen description-based predictor and response-share changes from to . The five-dose series changes sign between corruption fractions and . Pairing contrasts also appear on CUB and Flowers and under a second public initialisation, with a gate-limited result on ConvNeXt. A separate analysis relates geometric reallocation to perturbation-specific accuracy changes, whose associations vary across datasets. Attributing the direction to description content needs separable caption sets; a pre-training screen passes one of six caption pairs across five datasets, and the content comparison on that pair is not significant. Together, these results provide an affine-invariant diagnostic of pairing-dependent response allocation, leaving open whether description content determines its direction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.