Budget-constrained Sequential Active Learning to efficiently De-censor Survival Data
Abstract
Standard supervised learners attempt to learn a model from a labeled dataset. Given a small set of labeled instances, and a pool of unlabeled instances, a budgeted learner can use its given budget to pay to acquire the labels of some unlabeled instances, which it can then use to produce a model. Here, we explore budgeted learning in the context of survival datasets, which include (right) censored instances, where we know only a lower bound on an instance’s time-to-event. Here, that learner can pay to (partially) label a censored instance – e.g., to acquire the actual time for an instance [perhaps go from (3 yr, censored) to (7.2 yr, uncensored)], or other variants [e.g., learn about one more year, so go from (3 yr, censored) to either (4 yr, censored) or perhaps (3.2 yr, uncensored)]. This serves as a model of real world data collection, where follow-up with censored patients does not always lead to uncensoring, and how much information is given to the learner model during data collection is a function of the budget and the nature of the data itself. We provide both experimental and theoretical results for how to apply state-of-the- art budgeted learning algorithms to survival data and the respective limitations that exist in doing so. For an entropy-capped batch objective, marginal greedy selection under equal costs achieves the same (1 − 1/e) approximation factor as BatchBALD. Moreover, empirical analyses on several survival tasks show that our model performs better than other potential approaches on several benchmarks. Our acquisition rule, EPIG_surv, selects a patient and a follow-up depth by the expected reduction in predictive uncertainty on a target population per unit of expected cost. Probes are acquired in batches, with the model refitted between rounds.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.