OptAPI: Optimally Active Powered Inference under Label Budget
Abstract
Gold-standard labeled data are essential for reliable statistical inference but are often severely limited by annotation budgets. Existing active inference methods use pretrained predictive models to exploit unlabeled data; however, they typically assume all observations are equally important and optimize efficiency for only a single parameter dimension. To effectively capture joint multivariate uncertainty, we propose , a two-stage optimally active powered inference framework that maximizes the utility of both labeled and unlabeled data under a strict budget. The first stage deploys an active sampling rule to identify highly informative observations, while the second stage fine-tunes the prediction-powered correction. Crucially, both stages are unified under a D-optimality criterion that minimizes the determinant of the asymptotic covariance matrix, thereby directly targeting multidimensional uncertainty. Furthermore, we establish the asymptotic normality of the estimator, which theoretically justifies our D-optimal design and guarantees the validity of our constructed confidence intervals. Extensive numerical evaluations on synthetic and real-world datasets demonstrate the framework's superior effectiveness and statistical efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.