A Safe and Reliable Active Learning Method
Abstract
Active learning (AL) aims to alleviate the annotation cost of deep supervised learning. Recently, coverage-aware AL methods obtain state-of-the-art performance across a wide range of budget settings. However, existing approaches often pay limited attention to the purity component of the coverage-purity framework. Explicitly optimizing purity in AL is challenging because of the inherent trade-off between coverage and purity, as well as the lack of an effective mechanism for incorporating purity into the AL process. To address these issues, we introduce the safe zone framework which identifies reliable, high-purity regions within the unlabeled pool and thereby preserves the purity of the total covered area. We further provide theoretical guarantees for the purity achieved by the resulting safe zone. Subsequently, we propose SafeHerding, an efficient purity-aware AL method that prioritizes high-purity samples to improve model performance. The safe zone framework provides the foundation for key components of SafeHerding, including coverage optimization and kernel herding paradigm. In addition, SafeHerding incorporates a novel skipping strategy to manage the intrinsic trade-off between coverage and purity. Experimental results across different learning settings demonstrate that SafeHerding consistently outperforms existing baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.