Beyond Pseudo-Labeling: Retrieval-Augmented Visual Tuning for Semi-Supervised Learning in the Wild
Abstract
Semi-supervised learning (SSL) commonly exploits unlabeled data via confidence-thresholded pseudo-labeling, assuming that high confidence implies correctness. This assumption often fails in in-the-wild regimes with strong distribution mismatch, where thresholding induces two-sided selection bias (admitting high-confidence errors while discarding informative low-confidence samples) and shift further aggravates confidence miscalibration. Their coupling makes single-threshold heuristics brittle, leading to confirmation bias and even the "unlabeled hurts” phenomenon. We propose Retrieval-Augmented Visual Tuning (RAVT), which replaces pointwise pseudo-label acceptance with evidence-backed neighborhoods conditioned on labeled queries. RAVT constructs reliable neighborhoods through hybrid retrieval that combines global semantic recall with patch-level ranking, and strengthens semantic consistency using implicit textual cues; learning is driven by non-local neighborhood evidence and followed by supervised anchored tuning to stabilize representations. Across four progressively harder mismatch settings, RAVT is particularly effective in in-the-wild regimes, including cross-dataset and multi-source heterogeneous unlabeled pools, where it consistently mitigates negative transfer and delivers more stable cross-distribution generalization, achieving substantial gains over strong pseudo-labeling baselines; in standard near-i.i.d. SSL settings, RAVT can also serve as a plug-in component for existing pseudo-labeling methods and further improve their performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.