acceptodds
Under review as a conference paper at ICLR 2027

Learning Source Acquisition Policies by Offline Planning

Abstract

Prediction under an acquisition budget requires choosing which information to observe. A query that reveals a feature group spends resources that could support later queries, and its value can depend on the decisions it enables. O-MPAC transfers finite-horizon acquisition targets into a shared action-conditioned scorer. An offline teacher evaluates risk and source cost on complete training records; the learner uses partial observations and source metadata to predict acquisition values and preferences. Online execution re-scores sources after each query and enforces the available budget. We analyze how complete-record information, tied action targets and the remaining planning horizon affect this transfer. Uniform supervision over tied minima preserves the target distribution under source relabeling, while matching the final acquisition step separates informative and redundant queries through terminal risk. In a five-seed context-dependent routing experiment, this supervision achieves 0.965 accuracy under both the original and context-last source orders. On six real tasks, validation selects H1 noCE in all thirty splits. O-MPAC has the highest mean budget-integrated accuracy on five tasks against source-adapted GDFS, DIME, AACO+NN and a static policy. A separate native-pipeline comparison against GDFS and DIME gives the same five-task pattern.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.