Prior-Enlightened Positive-Unlabeled Learning
Abstract
Positive-Unlabeled (PU) learning trains models from limited labeled positive samples and unlabeled data with mixed positive and negative instances. A core challenge resides in extracting credible negative candidates from unlabeled data, a procedure that critically determines the ranking supervision quality of pairwise PU learning. Current sampling paradigms rely solely on intrinsic data statistics and real-time model outputs, while substantially neglecting informative external prior cues. In practical scenarios, such prior knowledge provides complementary evidence for identifying credible negative samples. This constitutes a critical research gap: whether and how external priors can be properly integrated to alleviate the intrinsic limitations of purely model-based sampling in PU learning? To tackle this issue, this work presents a Prior-Enlightened Positive-Unlabeled Learning (PEPU) framework for pairwise PU learning. PEPU integrates the external prior credibility of unlabeled negative candidates with model-derived pairwise utility to achieve adaptive sampling, and formulates a principled optimization objective that balances pairwise utility and prior knowledge. This work theoretically analyzes the trade-off between prior guidance and model optimization, and further investigates its effects on hidden positive contamination and downstream ranking performance. Extensive experiments on controlled PU studies and two real-data application domains validate that PEPU consistently outperforms state-of-the-art sampling strategies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.