acceptodds
Under review as a conference paper at ICLR 2027

RELATIONAL STRUCTURE RECOVERY FOR STRICT SCORE-ONLY VLM ADAPTATION

Abstract

Deployment interfaces for vision-language models may expose only a fixed tensor of prompt-level class scores, withholding images, embeddings, gradients, parameter updates, labels, and additional queries. Existing score-only methods mainly aggregate or reweight these scores. We ask instead whether the tensor contains a relational representation that supports target-conditioned adaptation. Prompt-class scores are repeated text-induced measurements: on non-degenerate rows, normalized class contrasts induce a shared geometry and form a maximal invariant to sample-wise offsets and positive scales, while prompt-dependent deviations provide complementary information after the shared response is accounted for. Prompt-Class Residual Interaction (PCRI) instantiates this shared–conditional decomposition by constructing a Score-space Core, extracting Core-conditioned prompt innovation, and using the joint representation only to correct the frozen semantic score field through target prototypes. With one configuration selected on a disjoint CIFAR-100 development setting and frozen across all evaluation domains, PCRI improves matched CARPRT by +2.378 percentage points on average across 11 datasets and three VLM backbones, with gains in 32/33 settings under identical cached scores. Matched controls attribute most Core-only gain to class-contrast standardization and isolate a further +0.306-point benefit from Core conditioning; support stress tests show that limited or unrepresentative target support can make the correction unreliable. PCRI uses no hidden representations or additional VLM calls after score extraction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.