Correcting Selection Bias in Positive and Unlabeled Generation with Selection Records
Abstract
Positive and unlabeled (PU) data combine labeled positives with a population sample whose classes are unknown. A generator fitted to the labeled positives inherits their selection bias: positives that are verified more often stay overrepresented, even when every label is correct. We correct this bias with selection records, which show whether objects in a population sample were selected for verification but not their classes. Under shared selection and constant reporting, the normalized density ratio of labeled positives to selected objects identifies the positive distribution without selection probabilities or the population frequency of positive labels. We bound the resulting distribution error by the ratio estimation error and guide a diffusion model with the fitted weights. Under simulated biased selection on MNIST, CelebA and CIFAR-10, the weights improve the group composition of the weighted positives, and joint class and group error falls on MNIST and CIFAR-10. On MNIST, generated composition also improves at matched fractions of predicted positives, although even weights equal to the true class leave generation error. On CelebA, longer ratio training lowers the joint error of the weighted target, but the held out classification loss does not select it. Further experiments separate the information in the records from their number and show that restricted probability models, mismatched records and class dependent selection prevent recovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.