Cross-Prompt Inference Mitigates Rare-Class Suppression in Imbalanced VLM Finetuning
Abstract
When a vision-language model is finetuned on class-imbalanced data, rare-class recall can fall well below the model’s own zero-shot performance. In this regime, we uncover a surprising matched-prompt penalty whereby the standard choice of reusing the descriptive finetuning prompt at inference can yield lower rare-class recall than using a different descriptive prompt. We call this inference-time strategy cross-prompt inference. We evaluate diverse VLMs across six datasets and find that cross-prompt inference consistently recovers rare-class recall, with varying trade-offs in frequent-class recall. The effect reflects a train-eval prompt interaction rather than prompt quality alone, and does not arise with non-descriptive prompt changes. It is strongest when imbalance and visual confusability produce substantial rare-class suppression, and is weak or absent when such suppression does not occur. Using Tuning Contribution (TuCo), which decomposes a model’s computation into pretrained and finetuning parts, we find that cross-prompt inference reduces reliance on finetuning-induced computation and shifts the model toward its pretrained component. Reducing that finetuning contribution also substantially reduces the cross-prompt advantage. These results show that rare-class knowledge suppressed by imbalanced finetuning can remain accessible under a different inference prompt.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.