Non-Myopic Active Feature Acquisition via Pathwise Policy Gradients
Abstract
Active feature acquisition (AFA) considers prediction problems in which features are costly to obtain and the learner adaptively decides which feature values to acquire for each instance and when to stop and predict. In this paper, we introduce a continuous relaxation of the acquisition process that enables non-myopic pathwise policy gradients (NM-PPG) through the full acquisition trajectory, avoiding the high variance of standard score-function policy gradients while allowing end-to-end optimization of the acquisition policy. To better align training with deployment, we develop a straight-through rollout that follows discrete feature acquisitions in the forward pass while backpropagating through the corresponding soft relaxation. We derive an AFA-specific average-case upper bound on the variance of this gradient estimator, which characterizes instability and motivates staged temperature sharpening. Experiments on both synthetic and real-world datasets demonstrate that NM-PPG yields superior performance relative to state-of-the-art AFA baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.