Koo-Fu CLIP: Closed-Form Adaptation of Vision–Language Models via Fukunaga–Koontz Linear Discriminant Analysis
Abstract
Vision-language models such as CLIP provide general-purpose representations, but their embeddings are not explicitly optimized for class separation or compactness in downstream supervised tasks. We propose Koo-Fu CLIP, a supervised adaptation method based on Fukunaga–Koontz Linear Discriminant Analysis. The method combines regularized within-class whitening with a rotation that orders projection directions by between-class scatter. The resulting closed-form transformation supports multiple embedding sizes through truncation, without iterative optimization. Experiments with CLIP, SigLIP2, and DINOv3 on seven datasets show improvements over the original embeddings in prototype and nearest-neighbor classification across most evaluated settings. Dimensionality reduction often further improves nearest-neighbor accuracy while reducing storage requirements and distance-computation costs. For example, on ImageNet-1k, Koo-Fu CLIP raises CLIP 15-nearest-neighbor accuracy from 79.42% to 82.18% while reducing embedding dimensionality from 768 to 192. Although learned linear projections generally achieve higher accuracy, Koo-Fu CLIP obtains all output dimensionalities from a single fitted transformation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.