WASSAL: Wasserstein Distance-Based Soft Subsetting for Active Learning
Abstract
We introduce Wasserstein Distance-Based Soft Subsetting for Active Learning (WASSAL), an approach that enhances active learning for image classification by explicitly using distributional information when selecting samples for expert labeling. Traditional AL methods often overlook this perspective and rely mainly on model uncertainty or geometric diversity. The key to our methodology is a Weighted Wasserstein adaptation formulation, which learns class-specific weights for the unlabeled samples so that the weighted samples are distributionally close to the labeled datapoints of a given class. WASSAL employs these learned weights in two ways: (i) samples with the least weights across classes are treated as most confusing from a distributional perspective and form the query set for expert labeling; (ii) samples with the highest weights are treated as typical/landmark examples of their respective classes and can be soft-labeled without expert consultation. We propose a novel weighted loss term involving these soft-labeled examples in the next training round. Hence WASSAL not only leverages expert knowledge on the few queried examples, but also utilizes information implicitly available in non-queried examples. Our formulation extends the Wasserstein core-set view of Mahmood et al. (2022) by making the matching class-conditional. Across five benchmarks, WASSAL and WASSALWITHSOFT lead every budget on Caltech-101 and STL-10, remain within a few points of the best baselines on CIFAR-10 and PneumoniaMNIST, and hold their own under the high-variance SVHN pool, showing that distributional soft subsetting is a practical addition to the active learning toolbox.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.