Calibration-Enhanced Approach to Biased Positive-Unlabeled Learning
Abstract
Positive-Unlabeled (PU) learning addresses the problem of training a binary classifier using an incomplete training dataset consisting of a subset of labeled positive instances and a set of unlabeled instances, which contains both positive and negative examples. A practically relevant and recently intensively studied scenario is biased PU learning, where the labeling mechanism determining which positive examples receive labels may depend on the feature values of the instances. A naive method in PU learning, known as the non-traditional classifier (NTC), treats all unlabeled instances as negative. However, this leads to a biased estimation of the true class posterior probability. More effective learning from biased PU data is possible but challenging, as it typically requires estimating the propensity score, defined as the instance-specific probability that a positive example is labeled. Existing propensity score estimation methods are often computationally expensive, since they rely on iterative procedures that alternately update the classifier and the propensity score model. In this work, we show that under relatively weak functional assumption on the propensity score, it is possible to accurately estimate the class posterior without explicitly modeling the former, provided that a small calibration set is available. The proposed approach (called PUCal) is based on an appropriate calibration of probabilities produced by the NTC on a small set of fully observable observations. We show that such a procedure leads to a more accurate classifier compared to existing methods for biased PU learning even when only a very limited number of observations are available in the calibration set. We also provide estimation error bounds for the proposed method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.