Do Shared VLM Coordinates Imply Shared Calibration? A Study through Few-Shot Class-Incremental Learning
Abstract
Pretrained vision-language models place visual and textual class references in shared coordinates, but alignment does not determine their compatibility with downstream geometric calibration. We expose this distinction through training-free few-shot class-incremental learning, where support-derived visual estimates and pretrained textual anchors are evaluated against the same image queries. Crossing visual- and text-derived diagonal calibration transforms with the two isolated scorers reveals a reproducible role-dependent calibration asymmetry. Across three encoders and four benchmarks, visual-derived calibration is more favorable to visual than textual scoring in all 12 settings. Textual accuracy decreases in every setting, even when the calibration uses textual statistics from base classes or all currently known classes. Our geometric analysis and source-matched counterexample explain the key distinction: calibration replaces native cosine with an affine Mahalanobis-like cosine, and matching reference marginals does not ensure preservation of image-to-reference class ordering. These results motivate source-preserving scoring: calibrate the empirical visual estimate, preserve the pretrained textual anchor, and compose their class evidence. We instantiate this principle as Dual-Space Product-of-Experts (DS-PoE), combining a base-calibrated visual expert and a native textual expert through fixed logit addition, without gradient updates, replay, or dataset-specific modality weights. Across four benchmarks, DS-PoE achieves state-of-the-art performance, attaining the best mean session accuracy on every benchmark and a four-benchmark mean of 78.6% under the evaluated training-free protocol. Our code is included in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.