Evidence Acquisition And Stopping Via Lattice Distillation
Abstract
Sequential evidence acquisition requires deciding what to observe and when to stop from partial information. We introduce a complete-record view of policy supervision: each fully observed training instance defines a lattice of evidence subsets, with a stop/query decision at every node. For a modest number of evidence groups, this structure can be enumerated even when individual groups contain rich representations. Our method, Subset Lattice Distillation (SLD), scores subsets with a shared masked predictor, constructs multi-step decision targets by cost-sensitive backward planning, and distills them into an observed-history controller. Uniformly resampled masks train the student across the subset space, while combinatorial planning remains offline. Separating subset scores from acquisition prices allows the predictor to be reused when the policy is replanned for new costs. Our analysis identifies action-margin conditions under which privileged supervision preserves optimal online decisions. On retrospective influenza monitoring, SLD achieves higher phase F1 at lower acquisition cost than both the closest-cost adaptive baselines and the fixed full-evidence model. A geographic and seasonal holdout, price perturbations, and ablation studies provide further evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.