acceptodds
Under review as a conference paper at ICLR 2027

Measure What Changes the Decision: Bounded Lookahead for Queried-Only Deployment

Abstract

When an action must be measured before it can be deployed, information acquisition and resource allocation become coupled. A measurement budget Q controls which action outcomes are known, while a separate budget B couples the measured actions selected for deployment. Under this coupling, a query can have little immediate value yet make a later query useful. Exact one-step Expected Deployable-Utility Gain (EDUG) cannot price this delayed payoff. We therefore introduce A2-EDUG, a bounded two-step policy that ranks all queries by their immediate decision value, evaluates two-query plans for the four best candidates using 256 posterior draws, executes only the selected first query, and replans. With the belief model, query budget, and deployment solver held fixed, A2-EDUG reduces average regret along the query path by 11.99% on fresh DIV2K (95% group-bootstrap interval [6.26%, 17.55%]), by 16.16% on a speaker-disjoint FSDD confirmation, and by 7.72% ([5.78%, 9.57%]) in a 240-episode FSDD scale study. In both systems, the policy is about 3% worse after its first query and becomes better after the second, matching the predicted delayed-payoff mechanism. A control matched to the same number of exact solver calls and a monotone k = 1, 2, 4, 8 candidate sweep show that the gain comes from two-step planning, not simply from running more one-step evaluations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.