Spectral Gradient Merging for Stable Few-Shot Adaptation of Vision-Language Models
Abstract
Few-shot adaptation of vision–language models remains sensitive to the learning rate and training duration, even when SGD achieves strong accuracy. We propose OLM-SGD, an optimizer that combines class-unique minibatching with orthogonal low-rank class-specific gradient merging. Each minibatch contains at most one example per class. For each parameter tensor, OLM-SGD retains dominant per-example singular components and merges them through polar reconstruction. It directly adapts the original encoder blocks without additional trainable modules or optimizer moments. Our analysis establishes retained-gradient alignment and energy control, with conditional stationarity guarantees and finite-shot population bounds that account for support-set reuse. Experiments span 11 datasets, three CLIP backbones, and five shot levels. OLM-SGD improves dataset-average accuracy over the strongest evaluated parameter-efficient or frozen-feature baseline in all 15 backbone–shot settings, with gains up to 1.9 percentage points. One-shot experiments over 300 epochs and four learning rates also show reduced learning-rate sensitivity and improved accuracy retention. At learning rate , the mean peak-to-final accuracy drop on ViT-B/16 decreases from 6.49 percentage points with SGD to 2.22 with OLM-SGD. Ablations show that polar reconstruction contributes beyond class-unique sampling and low-rank truncation alone. These results support class-specific gradient merging as an effective approach to accurate and stable few-shot adaptation. Our code is included in the supplementary materials and will be released upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.