acceptodds
Under review as a conference paper at ICLR 2027

Recovering Silenced Gradients: Noise-Robust Prompt Tuning for Vision-Language Models

Abstract

Prompt tuning efficiently adapts vision-language models by learning a small set of parameters while keeping the backbone frozen. However, existing noise-robust methods suppress useful learning signals along with corrupted supervision. Through theoretical and experimental analyses, we identify a failure mode termed gradient silencing. Correctly labeled samples with low prediction confidence, which we call hard positives, can be misclassified as noisy. Mean absolute error (MAE) then attenuates their corrective gradients in proportion to that confidence. To address this limitation, we propose Geometric Gradient Rectification Prompting (GGRP), which couples geometric target rectification with silenced gradient recovery. The key idea is to use image–text similarities to determine both training targets and sample contributions, rather than only separating clean and noisy samples. Specifically, temporal momentum guidance stabilizes textual class prototypes. We then use geometric target rectification to construct training targets via optimal transport (OT) alignment. Finally, silenced gradient recovery uses the OT assignment margin to balance cross-entropy on rectified targets with MAE on observed labels. This allows hard positives to receive corrective gradients even at low prediction confidence when OT sufficiently supports their correct labels, while retaining MAE-based attenuation for ambiguous samples. Extensive experiments on seven datasets with synthetic label noise and Food101N with naturally noisy labels show that GGRP surpasses state-of-the-art methods across a range of noise settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.