GSSP: Elevating Under-Performing Classes in Zero-Shot Image Classification
Abstract
Vision Language models exhibit strong generalized zero-shot image classification capabilities, however, they frequently suffer from severe performance disparities with accuracy in which under-performing classes often fall to near zero accuracy. To address this without the computational burden of retraining, we propose the Gram-Schmidt Subspace Projection (GSSP) a training-free, lightweight framework that operates directly in the post-projection embedding space. Our method constructs orthonormal bases using the corresponding text embeddings of the under-performing classes, applying targeted geometric transformations to the visual features. Because raw projection can cause catastrophic forgetting of a model's generalized knowledge, we introduce the extension GSSP-Boost, which utilizes a parametrized blending mechanism to regularize the method and maintain the overall performance. Through comprehensive empirical evaluations across a diverse suite of datasets, we demonstrate that GSSP-Boost significantly improves the under-performing classes' accuracy while preserving the model's overall performance. Notably, our approach consistently outperforms existing baselines across different architectures, offering a highly effective flexible method for deploying robust zero-shot image classification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.